# Replicate vs Modal

Canonical: https://inetgeek.com/compare/replicate-vs-modal/

Every value below is read from Replicate's and Modal's own documentation. See https://inetgeek.com/methodology/ for how.

## At a glance

| Criterion | Replicate | Modal |
| --- | --- | --- |
| Pricing model | Two meters, chosen by the model rather than by you. Public models bill per run — per output image, per thousand tokens, per second of video. Private models you deploy yourself run on dedicated hardware and bill for THE TIME THEY SPEND SETTING UP, THE TIME THEY SPEND IDLE and the time they spend active. Fast-booting fine-tunes are the exception and bill active time only. | A platform fee plus metered compute, and the fee buys a compute allowance rather than access. Starter is $0 a month and includes $30 of compute; Team is $250 a month and includes $100. Compute is billed per second on top of both. |
| H100 SXM, per GPU-hour | $5.49 per hour for a single Nvidia H100 with 80GB of GPU RAM, published as $0.001525 per second — the rate the per-second meter actually charges. | $3.95 per hour for an Nvidia H100 SXM5, published as $0.001097 per second. Modal's own worked example on the same page prices a fleet at $3.95 per GPU-hour. |
| Cheapest GPU, per hour | $0.81 per hour for an Nvidia T4 with 16GB, the cheapest GPU on the hardware list. A CPU-only instance is cheaper still at $0.09 an hour. | $0.59 per hour for an Nvidia T4, published as $0.000164 per second — the cheapest GPU on the resource table. |
| Billing granularity | Per second. Every hardware row publishes a per-second rate alongside the hourly one, and the per-second figure is the meter. | Per second, and per CPU core rather than per instance: a physical core is $0.0000131 per core per second with a minimum of 0.125 cores. |
| Support on entry paid plan | Nothing named on the standard product. A dedicated account manager, priority support, higher GPU limits and performance SLAs are all listed under Enterprise and volume discounts, reached by contacting Replicate rather than by picking a plan. | A community Slack on the lower plans; private Slack support and embedded ML engineering services are Enterprise. Support is a plan feature here rather than something every tier carries. |

## Where they differ

### Pricing model

- Replicate: Two meters, chosen by the model rather than by you. Public models bill per run — per output image, per thousand tokens, per second of video. Private models you deploy yourself run on dedicated hardware and bill for THE TIME THEY SPEND SETTING UP, THE TIME THEY SPEND IDLE and the time they spend active. Fast-booting fine-tunes are the exception and bill active time only. ([source](https://replicate.com/pricing))
- Modal: A platform fee plus metered compute, and the fee buys a compute allowance rather than access. Starter is $0 a month and includes $30 of compute; Team is $250 a month and includes $100. Compute is billed per second on top of both. ([source](https://modal.com/pricing))

### H100 SXM, per GPU-hour

- Replicate: $5.49 per hour for a single Nvidia H100 with 80GB of GPU RAM, published as $0.001525 per second — the rate the per-second meter actually charges. ([source](https://replicate.com/pricing))
- Modal: $3.95 per hour for an Nvidia H100 SXM5, published as $0.001097 per second. Modal's own worked example on the same page prices a fleet at $3.95 per GPU-hour. ([source](https://modal.com/pricing))

### Cheapest GPU, per hour

- Replicate: $0.81 per hour for an Nvidia T4 with 16GB, the cheapest GPU on the hardware list. A CPU-only instance is cheaper still at $0.09 an hour. ([source](https://replicate.com/pricing))
- Modal: $0.59 per hour for an Nvidia T4, published as $0.000164 per second — the cheapest GPU on the resource table. ([source](https://modal.com/pricing))

### Billing granularity

- Replicate: Per second. Every hardware row publishes a per-second rate alongside the hourly one, and the per-second figure is the meter. ([source](https://replicate.com/pricing))
- Modal: Per second, and per CPU core rather than per instance: a physical core is $0.0000131 per core per second with a minimum of 0.125 cores. ([source](https://modal.com/pricing))

### Support on entry paid plan

- Replicate: Nothing named on the standard product. A dedicated account manager, priority support, higher GPU limits and performance SLAs are all listed under Enterprise and volume discounts, reached by contacting Replicate rather than by picking a plan. ([source](https://replicate.com/pricing))
- Modal: A community Slack on the lower plans; private Slack support and embedded ML engineering services are Enterprise. Support is a plan feature here rather than something every tier carries. ([source](https://modal.com/pricing))

## Which should you choose?

Pick Replicate if Running someone else's open model without packaging anything. Thousands of public models bill per run — per output image, per thousand tokens, per second of video — so a workload that fires occasionally pays only for the firing, with no instance to keep warm.

Pick Modal if Bursty, self-deployed inference where the bill should follow the work. Compute meters per second and per CPU core — a physical core is $0.0000131 per core-second with a 0.125-core minimum — and the H100 hour is $3.95, the cheapest of the four here that publishes one.

Consider something else: Replicate — You are deploying your own model and it will sit idle. Private models bill the time they spend setting up and waiting as well as working; only fast-booting fine-tunes escape that. The H100 hour is $5.49, against $3.95 on Modal and $3.49 on RunPod.

Consider something else: Modal — You want a large GPU named with its memory before committing, or support without an Enterprise contract. The page lists B200 and B300 but states VRAM only for the A100 rows, so the ceiling is unpublished; SOC 2, HIPAA, RBAC and SSO are all Enterprise.

## Documented by only one

- Free tier: Modal: Yes
- Kind of free offer: Modal: Free credit — $30 of compute a month on Starter, recurring rather than one-off
- Entry paid plan: Modal: $0 a month on Starter — the plan carries no fee at all, and $30 of monthly compute is included before anything is billed. The first plan with a monthly fee is Team at $250.
- Scales to zero: Replicate: Not supported — Private models bill idle time; fast-booting fine-tunes do not
- Autoscaling: Replicate: Supported — Scales up and down automatically with traffic
- Largest GPU offered: Replicate: 160GB on the 2x Nvidia A100 (80GB) instance, the largest GPU RAM figure the hardware table states. Larger multi-GPU A100, H100 and H200 configurations are listed but carry no stated GPU RAM and are sold only with committed spend contracts.
- SOC 2: Modal: SOC 2 compliance is listed as an Enterprise-plan feature on the pricing page, alongside HIPAA compatibility, audit logs, RBAC and SSO. No report or certificate is linked from the page itself.
- HIPAA: Modal: HIPAA compatibility is listed as an Enterprise-plan feature, grouped with SOC 2 and audit logs. The page says compatibility rather than a signed BAA.

## Questions this comparison answers

**Should I choose Replicate or Modal?**

Pick Replicate if Running someone else's open model without packaging anything. Thousands of public models bill per run — per output image, per thousand tokens, per second of video — so a workload that fires occasionally pays only for the firing, with no instance to keep warm.
Pick Modal if Bursty, self-deployed inference where the bill should follow the work. Compute meters per second and per CPU core — a physical core is $0.0000131 per core-second with a 0.125-core minimum — and the H100 hour is $3.95, the cheapest of the four here that publishes one.
