Replicate vs RunPod
SEPT 2026 auditA comparison of Replicate and RunPod built from values read directly from each provider's own documentation, with the source recorded against every figure.
The short answer
Choose Replicate if…
Running someone else's open model without packaging anything. Thousands of public models bill per run — per output image, per thousand tokens, per second of video — so a workload that fires occasionally pays only for the firing, with no instance to keep warm.
Editorial · Palash Bagchi · approved
Choose RunPod if…
Short experiments and bursty inference — per-second billing and a $0.27 entry card mean an idle hour costs nothing and a small job is genuinely small.
Editorial · Palash Bagchi · approved
Consider something else if…
- Replicate: You are deploying your own model and it will sit idle. Private models bill the time they spend setting up and waiting as well as working; only fast-booting fine-tunes escape that. The H100 hour is $5.49, against $3.95 on Modal and $3.49 on RunPod.
- RunPod: The catalogue mixes community and secure capacity, so the cheapest quoted rate is not always the tier a production workload should sit on.
5 sourced criteria separate them — see where, with sources, below.
Pick the criteria you care about. The chart counts how many of them lean toward each provider — the same read as scanning the bars below, just totalled for the ones you chose.
Replicate 0
RunPod 0
At a glance.
| Criterion | Replicate | RunPod |
|---|---|---|
| Pricing | ||
| Pricing model | Two meters, chosen by the model rather than by you. Public models bill per run — per output image, per thousand tokens, per second of video. Private models you deploy yourself run on dedicated hardware and bill for THE TIME THEY SPEND SETTING UP, THE TIME THEY SPEND IDLE and the time they spend active. Fast-booting fine-tunes are the exception and bill active time only. | Prepaid credit balance drawn down by usage: all compute and storage billed per second, with no data transfer fees and no monthly plan tiers. |
| GPU | ||
| H100 SXM, per GPU-hour | $5.49 per hour for a single Nvidia H100 with 80GB of GPU RAM, published as $0.001525 per second — the rate the per-second meter actually charges. | $3.49 per hour for H100 SXM 80GB, single card. The PCIe variant is $2.89 and H100 NVL is $3.19. |
| Cheapest GPU, per hour | $0.81 per hour for an Nvidia T4 with 16GB, the cheapest GPU on the hardware list. A CPU-only instance is cheaper still at $0.09 an hour. | $0.27 per hour for an RTX A5000 with 24GB. RTX 3090 is $0.50, RTX 4090 $0.74, L40S $1.09. |
| Largest GPU offered | 160GB on the 2x Nvidia A100 (80GB) instance, the largest GPU RAM figure the hardware table states. Larger multi-GPU A100, H100 and H200 configurations are listed but carry no stated GPU RAM and are sold only with committed spend contracts. | 288GB per GPU on B300; B200 at 180GB is $6.79 an hour and H200 at 141GB is $4.59. |
| Billing granularity | Per second. Every hardware row publishes a per-second rate alongside the hourly one, and the per-second figure is the meter. | Per second, with per-hour rates shown for comparison. Clusters can be started with no commitment and scaled to 64 GPUs. |
Where they differ.
Pricing model
- Replicate
- Two meters, chosen by the model rather than by you. Public models bill per run — per output image, per thousand tokens, per second of video. Private models you deploy yourself run on dedicated hardware and bill for THE TIME THEY SPEND SETTING UP, THE TIME THEY SPEND IDLE and the time they spend active. Fast-booting fine-tunes are the exception and bill active time only.
- RunPod
- Prepaid credit balance drawn down by usage: all compute and storage billed per second, with no data transfer fees and no monthly plan tiers.
Sources (2) →Sources ↓
- Pricing – Replicate ↗
“e they spend setting up; the time they spend idle, waiting for requests; and the time they spend active, processing your requests. If you ge”
Read 2026-09-17 · official pricing
- Billing overview - Runpod Documentation ↗
“Runpod uses a credit-based billing system where you add funds to your account and charges are deducted as you use resources. All compute and storage charges are billed per second, with no fees for data transfer.”
Read 2026-09-13 · official docs
H100 SXM, per GPU-hour
- Replicate
- $5.49 per hour for a single Nvidia H100 with 80GB of GPU RAM, published as $0.001525 per second — the rate the per-second meter actually charges.
- RunPod
- $3.49 per hour for H100 SXM 80GB, single card. The PCIe variant is $2.89 and H100 NVL is $3.19.
Sources (2) →Sources ↓
- Pricing – Replicate ↗
“GPU RAM 160GB RAM 288GB Nvidia H100 GPU gpu-h100 $ 0.001525 /sec $ 5.49 /hr GPU 1x CPU 13x GPU RAM 80GB RAM 144GB Nvidia L40S GPU gpu-l40s”
Read 2026-09-17 · official pricing
- Pricing — RunPod ↗
“H100 SXM 80 GB VRAM 125 GB RAM 20 vCPUs $ 3.49 /hr [...] H100 PCIe 80 GB VRAM 188 GB RAM 16 vCPUs $ 2.89 /hr [...] H100 NVL 94 GB VRAM 94 GB RAM 16 vCPUs $ 3.19 /hr”
Read 2026-09-06 · official pricing
Cheapest GPU, per hour
- Replicate
- $0.81 per hour for an Nvidia T4 with 16GB, the cheapest GPU on the hardware list. A CPU-only instance is cheaper still at $0.09 an hour.
- RunPod
- $0.27 per hour for an RTX A5000 with 24GB. RTX 3090 is $0.50, RTX 4090 $0.74, L40S $1.09.
Sources (2) →Sources ↓
- Pricing – Replicate ↗
“x GPU RAM 96GB RAM 144GB Nvidia T4 GPU gpu-t4 $ 0.000225 /sec $ 0.81 /hr GPU 1x CPU 4x GPU RAM 16GB RAM 16GB Additional hardware 4x Nvidia A”
Read 2026-09-17 · official pricing
- Pricing — RunPod ↗
“RTX A5000 24 GB VRAM 25 GB RAM 9 vCPUs $ 0.27 /hr [...] RTX 3090 24 GB VRAM [...] $ 0.5 /hr [...] L40S 48 GB VRAM 94 GB RAM 16 vCPUs $ 1.09 /hr”
Read 2026-09-06 · official pricing
Largest GPU offered
160GB on the 2x Nvidia A100 (80GB) instance, the largest GPU RAM figure the hardware table states. Larger multi-GPU A100, H100 and H200 configurations are listed but carry no stated GPU RAM and are sold only with committed spend contracts.
288GB per GPU on B300; B200 at 180GB is $6.79 an hour and H200 at 141GB is $4.59.
Sources (2) →Sources ↓
- Pricing – Replicate ↗
“x GPU RAM 80GB RAM 144GB 2x Nvidia A100 (80GB) GPU gpu-a100-large-2x $ 0.002800 /sec $ 10.08 /hr GPU 2x CPU 20x GPU RAM 160GB RAM 288GB Nvid”
Read 2026-09-17 · official pricing
- Pricing — RunPod ↗
“B300 288 GB HBM3e 251 GB RAM 32 vCPUs [...] B200 180 GB VRAM 283 GB RAM 28 vCPUs $ 6.79 /hr [...] H200 141 GB VRAM 276 GB RAM 24 vCPUs $ 4.59 /hr”
Read 2026-09-06 · official pricing
Billing granularity
- Replicate
- Per second. Every hardware row publishes a per-second rate alongside the hourly one, and the per-second figure is the meter.
- RunPod
- Per second, with per-hour rates shown for comparison. Clusters can be started with no commitment and scaled to 64 GPUs.
Sources (2) →Sources ↓
- Pricing – Replicate ↗
“GPU RAM 160GB RAM 288GB Nvidia H100 GPU gpu-h100 $ 0.001525 /sec $ 5.49 /hr GPU 1x CPU 13x GPU RAM 80GB RAM 144GB Nvidia L40S GPU gpu-l40s”
Read 2026-09-17 · official pricing
- Pricing — RunPod ↗
“Per hour Per second [...] GPU clusters in minutes with no commitments. Scale up to 64 GPUs, attach shared storage, and pay only for what you use”
Read 2026-09-06 · official pricing
When these numbers change, hear about it. Sources are re-checked monthly; a repricing goes out as a short note.
Documented by only one.
These criteria are published by one provider and not the other. An absence here means we have not found a source, not that the feature is missing.
Questions this comparison answers.
Should I choose Replicate or RunPod?
Pick Replicate if Running someone else's open model without packaging anything. Thousands of public models bill per run — per output image, per thousand tokens, per second of video — so a workload that fires occasionally pays only for the firing, with no instance to keep warm.
Pick RunPod if Short experiments and bursty inference — per-second billing and a $0.27 entry card mean an idle hour costs nothing and a small job is genuinely small.
