Skip to content
inetGeek

Replicate vs Baseten

SEPT 2026 audit

A comparison of Replicate and Baseten built from values read directly from each provider's own documentation, with the source recorded against every figure.

Updated 30 criteria comparedSources last checked

The short answer

Choose Replicate if…

Running someone else's open model without packaging anything. Thousands of public models bill per run — per output image, per thousand tokens, per second of video — so a workload that fires occasionally pays only for the firing, with no instance to keep warm.

Editorial · Palash Bagchi · approved

Choose Baseten if…

Teams that want per-token APIs and dedicated GPUs from one vendor with no plan fee — minute-granularity billing means a model that runs for ten minutes costs ten minutes.

Editorial · Palash Bagchi · approved

Consider something else if…

  • Replicate: You are deploying your own model and it will sit idle. Private models bill the time they spend setting up and waiting as well as working; only fast-booting fine-tunes escape that. The H100 hour is $5.49, against $3.95 on Modal and $3.49 on RunPod.
  • Baseten: Its dedicated H100 works out at $6.50 an hour, roughly three times DeepInfra's, so steady GPU workloads pay a premium for the platform around them.

4 sourced criteria separate them — see where, with sources, below.

No affiliate links, sponsored placements or paid rankings appear on this site. Ordering follows the sourced data and the stated criteria.

01.

At a glance.

CriterionReplicateBaseten
Pricing
Pricing modelTwo meters, chosen by the model rather than by you. Public models bill per run — per output image, per thousand tokens, per second of video. Private models you deploy yourself run on dedicated hardware and bill for THE TIME THEY SPEND SETTING UP, THE TIME THEY SPEND IDLE and the time they spend active. Fast-booting fine-tunes are the exception and bill active time only.No platform fee: per-token Model APIs for open models, and dedicated deployments billed per minute of compute with volume discounts. The Pro tier buys priority access to high-demand GPUs rather than a lower rate.
GPU
H100 SXM, per GPU-hour
$5.49 per hour for a single Nvidia H100 with 80GB of GPU RAM, published as $0.001525 per second — the rate the per-second meter actually charges.
$6.50 per hour for a dedicated H100 80GB. Baseten publishes $0.10833 per MINUTE — the hourly figure is that times 60, ours rather than Baseten's, and the page offers an hourly toggle. A100 80GB is $0.06667 a minute and B200 $0.16633.
Largest GPU offered
160GB on the 2x Nvidia A100 (80GB) instance, the largest GPU RAM figure the hardware table states. Larger multi-GPU A100, H100 and H200 configurations are listed but carry no stated GPU RAM and are sold only with committed spend contracts.
180GB per GPU on a B200, at $0.16633 per minute — $9.98 an hour.
Support
Support on entry paid planNothing named on the standard product. A dedicated account manager, priority support, higher GPU limits and performance SLAs are all listed under Enterprise and volume discounts, reached by contacting Replicate rather than by picking a plan.Email and in-app chat support on Basic, which is pay-as-you-go from $0. Pro adds dedicated support on Slack and Zoom.

Only criteria both providers publish appear here; a tinted cell marks a real difference. Criteria only one of them documents are listed below, and an absence there means we found no source — not that the feature is missing. How we source this.

The small bar above a value is inetGeek's own lean toward that side — computed from the same facts shown, never a number the provider published. See the picker below "The short answer" to weigh only the criteria you care about.

02.

Where they differ.

Pricing model

Replicate
Two meters, chosen by the model rather than by you. Public models bill per run — per output image, per thousand tokens, per second of video. Private models you deploy yourself run on dedicated hardware and bill for THE TIME THEY SPEND SETTING UP, THE TIME THEY SPEND IDLE and the time they spend active. Fast-booting fine-tunes are the exception and bill active time only.
Baseten
No platform fee: per-token Model APIs for open models, and dedicated deployments billed per minute of compute with volume discounts. The Pro tier buys priority access to high-demand GPUs rather than a lower rate.
Sources (2) →
  • Pricing – Replicate ↗

    “e they spend setting up; the time they spend idle, waiting for requests; and the time they spend active, processing your requests. If you ge”

    Read 2026-09-17 · official pricing

  • Pricing — Baseten ↗

    “$0 per month, pay as you go [...] Only pay for the compute you use, down to the minute. Volume discounts available [...] Pro Unlimited autoscaling and priority compute access [...] Priority access to high-demand GPUs”

    Read 2026-09-06 · official pricing

H100 SXM, per GPU-hour

Replicate
$5.49 per hour for a single Nvidia H100 with 80GB of GPU RAM, published as $0.001525 per second — the rate the per-second meter actually charges.
Baseten
$6.50 per hour for a dedicated H100 80GB. Baseten publishes $0.10833 per MINUTE — the hourly figure is that times 60, ours rather than Baseten's, and the page offers an hourly toggle. A100 80GB is $0.06667 a minute and B200 $0.16633.
Sources (2) →
  • Pricing – Replicate ↗

    “GPU RAM 160GB RAM 288GB Nvidia H100 GPU gpu-h100 $ 0.001525 /sec $ 5.49 /hr GPU 1x CPU 13x GPU RAM 80GB RAM 144GB Nvidia L40S GPU gpu-l40s”

    Read 2026-09-17 · official pricing

  • Pricing — Baseten ↗

    “Dedicated Deployments Only pay for the compute you use, down to the minute. [...] Price per Minute Hour [...] A100 80 GiB VRAM $0.06667 [...] H100 80 GiB VRAM $0.10833 [...] B200 180 GiB VRAM $0.16633”

    Read 2026-09-06 · official pricing

Largest GPU offered

Replicate

160GB on the 2x Nvidia A100 (80GB) instance, the largest GPU RAM figure the hardware table states. Larger multi-GPU A100, H100 and H200 configurations are listed but carry no stated GPU RAM and are sold only with committed spend contracts.

Baseten

180GB per GPU on a B200, at $0.16633 per minute — $9.98 an hour.

Sources (2) →
  • Pricing – Replicate ↗

    “x GPU RAM 80GB RAM 144GB 2x Nvidia A100 (80GB) GPU gpu-a100-large-2x $ 0.002800 /sec $ 10.08 /hr GPU 2x CPU 20x GPU RAM 160GB RAM 288GB Nvid”

    Read 2026-09-17 · official pricing

  • Pricing — Baseten ↗

    “B200 180 GiB VRAM $0.16633 Deploy”

    Read 2026-09-06 · official pricing

Support on entry paid plan

Replicate
Nothing named on the standard product. A dedicated account manager, priority support, higher GPU limits and performance SLAs are all listed under Enterprise and volume discounts, reached by contacting Replicate rather than by picking a plan.
Baseten
Email and in-app chat support on Basic, which is pay-as-you-go from $0. Pro adds dedicated support on Slack and Zoom.
Sources (2) →
  • Pricing – Replicate ↗

    “uirements, we can offer: Dedicated account manager Priority support Higher GPU limits Performance SLAs Help with onboarding, custom models,”

    Read 2026-09-17 · official pricing

  • Baseten Pricing ↗

    “included in basic: dedicated deployments model apis training fast cold starts soc 2 type ii and hipaa compliant email and in-app chat support [...] hands-on engineering expertise dedicated support on slack and zoom”

    Read 2026-09-14 · official pricing

When these numbers change, hear about it. Sources are re-checked monthly; a repricing goes out as a short note.

Confirm by email; unsubscribe from any issue. Your address goes to Kit and nowhere else — what we do with it.

03.

Documented by only one.

These criteria are published by one provider and not the other. An absence here means we have not found a source, not that the feature is missing.

Free tierBaseten: No — New workspaces get starting credits for testing; usage beyond them is billed per minute.
Kind of free offerBaseten: Free credit
Entry paid planBaseten: $0 per month on the Basic plan — pay as you go, with dedicated deployments, model APIs and training included rather than gated behind a fee.
Scales to zeroReplicate: Not supported — Private models bill idle time; fast-booting fine-tunes do not
AutoscalingReplicate: Supported — Scales up and down automatically with traffic
RegionsBaseten: Two Baseten regions selectable for regional deployments: us (United States) and eu (European Union); other regions by request to support. GPU model deployments only; Chains, training jobs and shared CPU types cannot select a region.
Uptime SLABaseten: Custom SLAs on Enterprise. No percentage is published on the pricing page.
Input price, top modelBaseten: $1.40 per million input tokens for GLM-5.3, with cached input at $0.14 and output at $4.40. GLM-5.2 Fast, a speed variant, is $2.10.
Output price, top modelBaseten: $4.40 per million output tokens for GLM-5.3.
Input price, cheapest modelBaseten: $0.10 per million input tokens for GPT OSS 120B, with output at $0.50. GLM-5.3-Flash is $0.15 in and $0.50 out.
Context windowBaseten: 1,048K tokens on the DeepSeek V4 line and the GLM 5.2/5.3 family, the widest it serves; 262K on the Kimi K2 models, 200K on GLM 4.7 and 128K on GPT OSS 120B. Baseten's table publishes these in thousands, so the magnitude is 1,048,000 rather than the 1,048,576 a power-of-two reading would give — the vendor's own figure, not a conversion of it.
Max output tokensBaseten: 262K tokens on most models, and it is a separate ceiling from the context window rather than the remainder of it: DeepSeek V4 Flash 0731 allows 384K output inside a 1,048K window, while GLM 4.7, Nemotron Ultra and GPT OSS 120B cap output at their full context. The Inkling models are the outlier at 32K.
Prompt cachingBaseten: Supported
Rate limitsBaseten: Two limits, requests and tokens per minute, set by account status rather than spend: an unverified Basic account gets 15 RPM and 100,000 TPM, a verified Basic or Pro account 120 RPM and 500,000 TPM, Enterprise custom. Cached input counts toward TPM at full weight even though it is billed cheaper.
Cheapest GPU, per hourReplicate: $0.81 per hour for an Nvidia T4 with 16GB, the cheapest GPU on the hardware list. A CPU-only instance is cheaper still at $0.09 an hour.
Billing granularityReplicate: Per second. Every hardware row publishes a per-second rate alongside the hourly one, and the per-second figure is the meter.
SOC 2Baseten: Maintains SOC 2 Type II certification for the inference platform; policies and certifications are held in the Baseten Trust Center.
ISO 27001Baseten: ISO 27001:2022 listed as a compliance program; an ISO 27001 certificate for Baseten Labs, Inc. is published in the Trust Center for request. Certificate is request-access in the Trust Center, not a public download.
HIPAABaseten: Maintains HIPAA compliance alongside SOC 2 Type II for model inference on the Baseten platform.
GDPR / data residencyBaseten: Baseten Cloud supports a GDPR compliance program; residency is set with regional environments or a region on a single deployment.
PCI DSSBaseten: PCI DSS - SAQ D listed as a compliance program; a PCI-DSS v4.0.1 AOC for SAQ D Service Provider is published in the Trust Center for request. SAQ D self-assessment AOC, not a QSA Report on Compliance.
Encryption at restBaseten: Trust Center control: datastores housing sensitive customer data are encrypted at rest, and transmission over public networks is encrypted.
Customer-managed keysBaseten: Runtime OIDC BYOK recipes: a model fetches customer-managed keys at runtime to decrypt encrypted weights and to decrypt requests and encrypt responses. Documented under runtime OIDC use cases, implemented in model code via the truss-examples envelope-encryption recipes.
Role-based access controlBaseten: Role-based access control with three organization roles - Admin, Member and Viewer - in a single-team organization.
Private networking / VPCBaseten: Single-tenant Enterprise environments expose endpoints via AWS PrivateLink or Google Cloud Private Service Connect, off the public internet.
Dedicated infrastructureBaseten: Single-tenant runs workloads in an isolated VPC in Baseten Cloud with compute restricted to your organization; Enterprise, custom pricing.
04.

Questions this comparison answers.

Should I choose Replicate or Baseten?

Pick Replicate if Running someone else's open model without packaging anything. Thousands of public models bill per run — per output image, per thousand tokens, per second of video — so a workload that fires occasionally pays only for the firing, with no instance to keep warm.

Pick Baseten if Teams that want per-token APIs and dedicated GPUs from one vendor with no plan fee — minute-granularity billing means a model that runs for ten minutes costs ten minutes.

Provider pages

Related comparisons

Every page here is sourced and dated.