Replicate vs Modal
SEPT 2026 auditA comparison of Replicate and Modal built from values read directly from each provider's own documentation, with the source recorded against every figure.
The short answer
Choose Replicate if…
Running someone else's open model without packaging anything. Thousands of public models bill per run — per output image, per thousand tokens, per second of video — so a workload that fires occasionally pays only for the firing, with no instance to keep warm.
Editorial · Palash Bagchi · approved
Choose Modal if…
Bursty, self-deployed inference where the bill should follow the work. Compute meters per second and per CPU core — a physical core is $0.0000131 per core-second with a 0.125-core minimum — and the H100 hour is $3.95, the cheapest of the four here that publishes one.
Editorial · Palash Bagchi · approved
Consider something else if…
- Replicate: You are deploying your own model and it will sit idle. Private models bill the time they spend setting up and waiting as well as working; only fast-booting fine-tunes escape that. The H100 hour is $5.49, against $3.95 on Modal and $3.49 on RunPod.
- Modal: You want a large GPU named with its memory before committing, or support without an Enterprise contract. The page lists B200 and B300 but states VRAM only for the A100 rows, so the ceiling is unpublished; SOC 2, HIPAA, RBAC and SSO are all Enterprise.
5 sourced criteria separate them — see where, with sources, below.
Pick the criteria you care about. The chart counts how many of them lean toward each provider — the same read as scanning the bars below, just totalled for the ones you chose.
Replicate 0
Modal 0
At a glance.
| Criterion | Replicate | Modal |
|---|---|---|
| Pricing | ||
| Pricing model | Two meters, chosen by the model rather than by you. Public models bill per run — per output image, per thousand tokens, per second of video. Private models you deploy yourself run on dedicated hardware and bill for THE TIME THEY SPEND SETTING UP, THE TIME THEY SPEND IDLE and the time they spend active. Fast-booting fine-tunes are the exception and bill active time only. | A platform fee plus metered compute, and the fee buys a compute allowance rather than access. Starter is $0 a month and includes $30 of compute; Team is $250 a month and includes $100. Compute is billed per second on top of both. |
| GPU | ||
| H100 SXM, per GPU-hour | $5.49 per hour for a single Nvidia H100 with 80GB of GPU RAM, published as $0.001525 per second — the rate the per-second meter actually charges. | $3.95 per hour for an Nvidia H100 SXM5, published as $0.001097 per second. Modal's own worked example on the same page prices a fleet at $3.95 per GPU-hour. |
| Cheapest GPU, per hour | $0.81 per hour for an Nvidia T4 with 16GB, the cheapest GPU on the hardware list. A CPU-only instance is cheaper still at $0.09 an hour. | $0.59 per hour for an Nvidia T4, published as $0.000164 per second — the cheapest GPU on the resource table. |
| Billing granularity | Per second. Every hardware row publishes a per-second rate alongside the hourly one, and the per-second figure is the meter. | Per second, and per CPU core rather than per instance: a physical core is $0.0000131 per core per second with a minimum of 0.125 cores. |
| Support | ||
| Support on entry paid plan | Nothing named on the standard product. A dedicated account manager, priority support, higher GPU limits and performance SLAs are all listed under Enterprise and volume discounts, reached by contacting Replicate rather than by picking a plan. | A community Slack on the lower plans; private Slack support and embedded ML engineering services are Enterprise. Support is a plan feature here rather than something every tier carries. |
Where they differ.
Pricing model
- Replicate
- Two meters, chosen by the model rather than by you. Public models bill per run — per output image, per thousand tokens, per second of video. Private models you deploy yourself run on dedicated hardware and bill for THE TIME THEY SPEND SETTING UP, THE TIME THEY SPEND IDLE and the time they spend active. Fast-booting fine-tunes are the exception and bill active time only.
- Modal
- A platform fee plus metered compute, and the fee buys a compute allowance rather than access. Starter is $0 a month and includes $30 of compute; Team is $250 a month and includes $100. Compute is billed per second on top of both.
Sources (2) →Sources ↓
- Pricing – Replicate ↗
“e they spend setting up; the time they spend idle, waiting for requests; and the time they spend active, processing your requests. If you ge”
Read 2026-09-17 · official pricing
- Plans and Pricing — Modal ↗
“/ mo free PRICING PLANS Starter $0 + compute / month Built for small teams and independent developers looking to level up. Get started with”
Read 2026-09-17 · official pricing
H100 SXM, per GPU-hour
- Replicate
- $5.49 per hour for a single Nvidia H100 with 80GB of GPU RAM, published as $0.001525 per second — the rate the per-second meter actually charges.
- Modal
- $3.95 per hour for an Nvidia H100 SXM5, published as $0.001097 per second. Modal's own worked example on the same page prices a fleet at $3.95 per GPU-hour.
Sources (2) →Sources ↓
- Pricing – Replicate ↗
“GPU RAM 160GB RAM 288GB Nvidia H100 GPU gpu-h100 $ 0.001525 /sec $ 5.49 /hr GPU 1x CPU 13x GPU RAM 80GB RAM 144GB Nvidia L40S GPU gpu-l40s”
Read 2026-09-17 · official pricing
- Plans and Pricing — Modal ↗
“H200 SXM $0.001261 / sec Nvidia H100 SXM5 $0.001097 / sec Nvidia RTX PRO 6000 $0.000842 / sec Nvidia A100, 80 GB $0.000694 / sec Nvidia A100”
Read 2026-09-17 · official pricing
Cheapest GPU, per hour
- Replicate
- $0.81 per hour for an Nvidia T4 with 16GB, the cheapest GPU on the hardware list. A CPU-only instance is cheaper still at $0.09 an hour.
- Modal
- $0.59 per hour for an Nvidia T4, published as $0.000164 per second — the cheapest GPU on the resource table.
Sources (2) →Sources ↓
- Pricing – Replicate ↗
“x GPU RAM 96GB RAM 144GB Nvidia T4 GPU gpu-t4 $ 0.000225 /sec $ 0.81 /hr GPU 1x CPU 4x GPU RAM 16GB RAM 16GB Additional hardware 4x Nvidia A”
Read 2026-09-17 · official pricing
- Plans and Pricing — Modal ↗
“vidia L4 $0.000222 / sec Nvidia T4 $0.000164 / sec CPU Physical core (2 vCPU equivalent ) $0.0000131 / core / sec *minimum of 0.125 cores pe”
Read 2026-09-17 · official pricing
Billing granularity
- Replicate
- Per second. Every hardware row publishes a per-second rate alongside the hourly one, and the per-second figure is the meter.
- Modal
- Per second, and per CPU core rather than per instance: a physical core is $0.0000131 per core per second with a minimum of 0.125 cores.
Sources (2) →Sources ↓
- Pricing – Replicate ↗
“GPU RAM 160GB RAM 288GB Nvidia H100 GPU gpu-h100 $ 0.001525 /sec $ 5.49 /hr GPU 1x CPU 13x GPU RAM 80GB RAM 144GB Nvidia L40S GPU gpu-l40s”
Read 2026-09-17 · official pricing
- Plans and Pricing — Modal ↗
“vidia L4 $0.000222 / sec Nvidia T4 $0.000164 / sec CPU Physical core (2 vCPU equivalent ) $0.0000131 / core / sec *minimum of 0.125 cores pe”
Read 2026-09-17 · official pricing
Support on entry paid plan
- Replicate
- Nothing named on the standard product. A dedicated account manager, priority support, higher GPU limits and performance SLAs are all listed under Enterprise and volume discounts, reached by contacting Replicate rather than by picking a plan.
- Modal
- A community Slack on the lower plans; private Slack support and embedded ML engineering services are Enterprise. Support is a plan feature here rather than something every tier carries.
Sources (2) →Sources ↓
- Pricing – Replicate ↗
“uirements, we can offer: Dedicated account manager Priority support Higher GPU limits Performance SLAs Help with onboarding, custom models,”
Read 2026-09-17 · official pricing
- Plans and Pricing — Modal ↗
“nvironment-level budgets Support via private Slack Audit logs, SAML SSO, and HIPAA Credit grants for startups Early-stage startups can get f”
Read 2026-09-17 · official pricing
When these numbers change, hear about it. Sources are re-checked monthly; a repricing goes out as a short note.
Documented by only one.
These criteria are published by one provider and not the other. An absence here means we have not found a source, not that the feature is missing.
Questions this comparison answers.
Should I choose Replicate or Modal?
Pick Replicate if Running someone else's open model without packaging anything. Thousands of public models bill per run — per output image, per thousand tokens, per second of video — so a workload that fires occasionally pays only for the firing, with no instance to keep warm.
Pick Modal if Bursty, self-deployed inference where the bill should follow the work. Compute meters per second and per CPU core — a physical core is $0.0000131 per core-second with a 0.125-core minimum — and the H100 hour is $3.95, the cheapest of the four here that publishes one.
