Replicate vs Baseten
SEPT 2026 auditA comparison of Replicate and Baseten built from values read directly from each provider's own documentation, with the source recorded against every figure.
The short answer
Choose Replicate if…
Running someone else's open model without packaging anything. Thousands of public models bill per run — per output image, per thousand tokens, per second of video — so a workload that fires occasionally pays only for the firing, with no instance to keep warm.
Editorial · Palash Bagchi · approved
Choose Baseten if…
Teams that want per-token APIs and dedicated GPUs from one vendor with no plan fee — minute-granularity billing means a model that runs for ten minutes costs ten minutes.
Editorial · Palash Bagchi · approved
Consider something else if…
- Replicate: You are deploying your own model and it will sit idle. Private models bill the time they spend setting up and waiting as well as working; only fast-booting fine-tunes escape that. The H100 hour is $5.49, against $3.95 on Modal and $3.49 on RunPod.
- Baseten: Its dedicated H100 works out at $6.50 an hour, roughly three times DeepInfra's, so steady GPU workloads pay a premium for the platform around them.
4 sourced criteria separate them — see where, with sources, below.
Pick the criteria you care about. The chart counts how many of them lean toward each provider — the same read as scanning the bars below, just totalled for the ones you chose.
Replicate 0
Baseten 0
At a glance.
| Criterion | Replicate | Baseten |
|---|---|---|
| Pricing | ||
| Pricing model | Two meters, chosen by the model rather than by you. Public models bill per run — per output image, per thousand tokens, per second of video. Private models you deploy yourself run on dedicated hardware and bill for THE TIME THEY SPEND SETTING UP, THE TIME THEY SPEND IDLE and the time they spend active. Fast-booting fine-tunes are the exception and bill active time only. | No platform fee: per-token Model APIs for open models, and dedicated deployments billed per minute of compute with volume discounts. The Pro tier buys priority access to high-demand GPUs rather than a lower rate. |
| GPU | ||
| H100 SXM, per GPU-hour | $5.49 per hour for a single Nvidia H100 with 80GB of GPU RAM, published as $0.001525 per second — the rate the per-second meter actually charges. | $6.50 per hour for a dedicated H100 80GB. Baseten publishes $0.10833 per MINUTE — the hourly figure is that times 60, ours rather than Baseten's, and the page offers an hourly toggle. A100 80GB is $0.06667 a minute and B200 $0.16633. |
| Largest GPU offered | 160GB on the 2x Nvidia A100 (80GB) instance, the largest GPU RAM figure the hardware table states. Larger multi-GPU A100, H100 and H200 configurations are listed but carry no stated GPU RAM and are sold only with committed spend contracts. | 180GB per GPU on a B200, at $0.16633 per minute — $9.98 an hour. |
| Support | ||
| Support on entry paid plan | Nothing named on the standard product. A dedicated account manager, priority support, higher GPU limits and performance SLAs are all listed under Enterprise and volume discounts, reached by contacting Replicate rather than by picking a plan. | Email and in-app chat support on Basic, which is pay-as-you-go from $0. Pro adds dedicated support on Slack and Zoom. |
Where they differ.
Pricing model
- Replicate
- Two meters, chosen by the model rather than by you. Public models bill per run — per output image, per thousand tokens, per second of video. Private models you deploy yourself run on dedicated hardware and bill for THE TIME THEY SPEND SETTING UP, THE TIME THEY SPEND IDLE and the time they spend active. Fast-booting fine-tunes are the exception and bill active time only.
- Baseten
- No platform fee: per-token Model APIs for open models, and dedicated deployments billed per minute of compute with volume discounts. The Pro tier buys priority access to high-demand GPUs rather than a lower rate.
Sources (2) →Sources ↓
- Pricing – Replicate ↗
“e they spend setting up; the time they spend idle, waiting for requests; and the time they spend active, processing your requests. If you ge”
Read 2026-09-17 · official pricing
- Pricing — Baseten ↗
“$0 per month, pay as you go [...] Only pay for the compute you use, down to the minute. Volume discounts available [...] Pro Unlimited autoscaling and priority compute access [...] Priority access to high-demand GPUs”
Read 2026-09-06 · official pricing
H100 SXM, per GPU-hour
- Replicate
- $5.49 per hour for a single Nvidia H100 with 80GB of GPU RAM, published as $0.001525 per second — the rate the per-second meter actually charges.
- Baseten
- $6.50 per hour for a dedicated H100 80GB. Baseten publishes $0.10833 per MINUTE — the hourly figure is that times 60, ours rather than Baseten's, and the page offers an hourly toggle. A100 80GB is $0.06667 a minute and B200 $0.16633.
Sources (2) →Sources ↓
- Pricing – Replicate ↗
“GPU RAM 160GB RAM 288GB Nvidia H100 GPU gpu-h100 $ 0.001525 /sec $ 5.49 /hr GPU 1x CPU 13x GPU RAM 80GB RAM 144GB Nvidia L40S GPU gpu-l40s”
Read 2026-09-17 · official pricing
- Pricing — Baseten ↗
“Dedicated Deployments Only pay for the compute you use, down to the minute. [...] Price per Minute Hour [...] A100 80 GiB VRAM $0.06667 [...] H100 80 GiB VRAM $0.10833 [...] B200 180 GiB VRAM $0.16633”
Read 2026-09-06 · official pricing
Largest GPU offered
160GB on the 2x Nvidia A100 (80GB) instance, the largest GPU RAM figure the hardware table states. Larger multi-GPU A100, H100 and H200 configurations are listed but carry no stated GPU RAM and are sold only with committed spend contracts.
180GB per GPU on a B200, at $0.16633 per minute — $9.98 an hour.
Sources (2) →Sources ↓
- Pricing – Replicate ↗
“x GPU RAM 80GB RAM 144GB 2x Nvidia A100 (80GB) GPU gpu-a100-large-2x $ 0.002800 /sec $ 10.08 /hr GPU 2x CPU 20x GPU RAM 160GB RAM 288GB Nvid”
Read 2026-09-17 · official pricing
- Pricing — Baseten ↗
“B200 180 GiB VRAM $0.16633 Deploy”
Read 2026-09-06 · official pricing
Support on entry paid plan
- Replicate
- Nothing named on the standard product. A dedicated account manager, priority support, higher GPU limits and performance SLAs are all listed under Enterprise and volume discounts, reached by contacting Replicate rather than by picking a plan.
- Baseten
- Email and in-app chat support on Basic, which is pay-as-you-go from $0. Pro adds dedicated support on Slack and Zoom.
Sources (2) →Sources ↓
- Pricing – Replicate ↗
“uirements, we can offer: Dedicated account manager Priority support Higher GPU limits Performance SLAs Help with onboarding, custom models,”
Read 2026-09-17 · official pricing
- Baseten Pricing ↗
“included in basic: dedicated deployments model apis training fast cold starts soc 2 type ii and hipaa compliant email and in-app chat support [...] hands-on engineering expertise dedicated support on slack and zoom”
Read 2026-09-14 · official pricing
When these numbers change, hear about it. Sources are re-checked monthly; a repricing goes out as a short note.
Documented by only one.
These criteria are published by one provider and not the other. An absence here means we have not found a source, not that the feature is missing.
Questions this comparison answers.
Should I choose Replicate or Baseten?
Pick Replicate if Running someone else's open model without packaging anything. Thousands of public models bill per run — per output image, per thousand tokens, per second of video — so a workload that fires occasionally pays only for the firing, with no instance to keep warm.
Pick Baseten if Teams that want per-token APIs and dedicated GPUs from one vendor with no plan fee — minute-granularity billing means a model that runs for ten minutes costs ten minutes.
