# RunPod vs Baseten

Canonical: https://inetgeek.com/compare/runpod-vs-baseten/

Every value below is read from RunPod's and Baseten's own documentation. See https://inetgeek.com/methodology/ for how.

## At a glance

| Criterion | RunPod | Baseten |
| --- | --- | --- |
| Entry paid plan | No plan tiers: a new account can start with as little as $10 in prepaid credits, and deploying a Pod requires at least one hour of credits for the chosen configuration. | $0 per month on the Basic plan — pay as you go, with dedicated deployments, model APIs and training included rather than gated behind a fee. |
| Pricing model | Prepaid credit balance drawn down by usage: all compute and storage billed per second, with no data transfer fees and no monthly plan tiers. | No platform fee: per-token Model APIs for open models, and dedicated deployments billed per minute of compute with volume discounts. The Pro tier buys priority access to high-demand GPUs rather than a lower rate. |
| Regions | 31 regions across the US, Europe, Asia and Australia for on-demand GPU instances. | Two Baseten regions selectable for regional deployments: us (United States) and eu (European Union); other regions by request to support. GPU model deployments only; Chains, training jobs and shared CPU types cannot select a region. |
| Uptime SLA | SLA-backed uptime is offered on reserved clusters, which are sold by contract. No percentage is published on the pricing page, and on-demand pods carry no stated commitment. | Custom SLAs on Enterprise. No percentage is published on the pricing page. |
| H100 SXM, per GPU-hour | $3.49 per hour for H100 SXM 80GB, single card. The PCIe variant is $2.89 and H100 NVL is $3.19. | $6.50 per hour for a dedicated H100 80GB. Baseten publishes $0.10833 per MINUTE — the hourly figure is that times 60, ours rather than Baseten's, and the page offers an hourly toggle. A100 80GB is $0.06667 a minute and B200 $0.16633. |
| Largest GPU offered | 288GB per GPU on B300; B200 at 180GB is $6.79 an hour and H200 at 141GB is $4.59. | 180GB per GPU on a B200, at $0.16633 per minute — $9.98 an hour. |
| SOC 2 | SOC 2 Type II and SOC 3 examinations completed; reports are listed in the Runpod Trust Center and some require approved access before download. | Maintains SOC 2 Type II certification for the inference platform; policies and certifications are held in the Baseten Trust Center. |
| ISO 27001 | ISMS certified to ISO/IEC 27001:2022 by Sensiba LLP; certificate RUN-ISMS-20260827 valid 27 August 2026 through 26 August 2029. Scope limited to the ISMS supporting Runpod's services and platform, not individual products. | ISO 27001:2022 listed as a compliance program; an ISO 27001 certificate for Baseten Labs, Inc. is published in the Trust Center for request. Certificate is request-access in the Trust Center, not a public download. |
| HIPAA | Maintains a HIPAA program; HIPAA sits among the Trust Center compliance resources, some of which require approved access before download. | Maintains HIPAA compliance alongside SOC 2 Type II for model inference on the Baseten platform. |
| GDPR / data residency | Stated fully GDPR compliant for data processed in European data center regions, with documented EU procedures for collection, storage, processing and deletion. | Baseten Cloud supports a GDPR compliance program; residency is set with regional environments or a region on a single deployment. |
| PCI DSS | Claimed for Secure Cloud's vetted infrastructure partners, who meet enterprise standards including SOC 2, ISO 27001 and PCI DSS certifications. Stated of the data-centre partners behind Secure Cloud, not of Runpod Inc. itself. | PCI DSS - SAQ D listed as a compliance program; a PCI-DSS v4.0.1 AOC for SAQ D Service Provider is published in the Trust Center for request. SAQ D self-assessment AOC, not a QSA Report on Compliance. |
| Encryption at rest | Optional per-Pod volume encryption: the volume disk is encrypted at rest on the host machine, but container disk and network volumes cannot be encrypted. | Trust Center control: datastores housing sensitive customer data are encrypted at rest, and transmission over public networks is encrypted. |
| Customer-managed keys | Not supported: Runpod stores the volume encryption key, which cannot be retrieved, and the docs state bring your own key is not supported. | Runtime OIDC BYOK recipes: a model fetches customer-managed keys at runtime to decrypt encrypted weights and to decrypt requests and encrypt responses. Documented under runtime OIDC use cases, implemented in model code via the truss-examples envelope-encryption recipes. |
| Role-based access control | Four team roles — Basic, Billing, Dev and Admin — each with set permissions; Admin has unrestricted access to members, settings, billing and resources. | Role-based access control with three organization roles - Admin, Member and Viewer - in a single-team organization. |
| Private networking / VPC | Global networking gives each Pod a private IP reachable only by other Pods in the same account; NVIDIA GPU Pods only, in 17 data centers. | Single-tenant Enterprise environments expose endpoints via AWS PrivateLink or Google Cloud Private Service Connect, off the public internet. |
| Dedicated infrastructure | Reserved Clusters: dedicated GPU clusters with guaranteed availability, custom configurations and SLA-backed uptime for enterprises scaling to 10,000+ GPUs. | Single-tenant runs workloads in an isolated VPC in Baseten Cloud with compute restricted to your organization; Enterprise, custom pricing. |

## Where they differ

### Entry paid plan

- RunPod: No plan tiers: a new account can start with as little as $10 in prepaid credits, and deploying a Pod requires at least one hour of credits for the chosen configuration. ([source](https://docs.runpod.io/accounts-billing/billing))
- Baseten: $0 per month on the Basic plan — pay as you go, with dedicated deployments, model APIs and training included rather than gated behind a fee. ([source](https://www.baseten.co/pricing/))

### Pricing model

- RunPod: Prepaid credit balance drawn down by usage: all compute and storage billed per second, with no data transfer fees and no monthly plan tiers. ([source](https://docs.runpod.io/accounts-billing/billing))
- Baseten: No platform fee: per-token Model APIs for open models, and dedicated deployments billed per minute of compute with volume discounts. The Pro tier buys priority access to high-demand GPUs rather than a lower rate. ([source](https://www.baseten.co/pricing/))

### Regions

- RunPod: 31 regions across the US, Europe, Asia and Australia for on-demand GPU instances. ([source](https://www.runpod.io/product/cloud-gpus))
- Baseten: Two Baseten regions selectable for regional deployments: us (United States) and eu (European Union); other regions by request to support. GPU model deployments only; Chains, training jobs and shared CPU types cannot select a region. ([source](https://docs.baseten.co/deployment/regional-deployments))

### Uptime SLA

- RunPod: SLA-backed uptime is offered on reserved clusters, which are sold by contract. No percentage is published on the pricing page, and on-demand pods carry no stated commitment. ([source](https://www.runpod.io/pricing))
- Baseten: Custom SLAs on Enterprise. No percentage is published on the pricing page. ([source](https://www.baseten.co/pricing/))

### H100 SXM, per GPU-hour

- RunPod: $3.49 per hour for H100 SXM 80GB, single card. The PCIe variant is $2.89 and H100 NVL is $3.19. ([source](https://www.runpod.io/pricing))
- Baseten: $6.50 per hour for a dedicated H100 80GB. Baseten publishes $0.10833 per MINUTE — the hourly figure is that times 60, ours rather than Baseten's, and the page offers an hourly toggle. A100 80GB is $0.06667 a minute and B200 $0.16633. ([source](https://www.baseten.co/pricing/))

### Largest GPU offered

- RunPod: 288GB per GPU on B300; B200 at 180GB is $6.79 an hour and H200 at 141GB is $4.59. ([source](https://www.runpod.io/pricing))
- Baseten: 180GB per GPU on a B200, at $0.16633 per minute — $9.98 an hour. ([source](https://www.baseten.co/pricing/))

### SOC 2

- RunPod: SOC 2 Type II and SOC 3 examinations completed; reports are listed in the Runpod Trust Center and some require approved access before download. ([source](https://www.runpod.io/legal/compliance))
- Baseten: Maintains SOC 2 Type II certification for the inference platform; policies and certifications are held in the Baseten Trust Center. ([source](https://docs.baseten.co/observability/security))

### ISO 27001

- RunPod: ISMS certified to ISO/IEC 27001:2022 by Sensiba LLP; certificate RUN-ISMS-20260827 valid 27 August 2026 through 26 August 2029. Scope limited to the ISMS supporting Runpod's services and platform, not individual products. ([source](https://www.runpod.io/legal/compliance))
- Baseten: ISO 27001:2022 listed as a compliance program; an ISO 27001 certificate for Baseten Labs, Inc. is published in the Trust Center for request. Certificate is request-access in the Trust Center, not a public download. ([source](https://trust.baseten.co/))

### HIPAA

- RunPod: Maintains a HIPAA program; HIPAA sits among the Trust Center compliance resources, some of which require approved access before download. ([source](https://www.runpod.io/legal/compliance))
- Baseten: Maintains HIPAA compliance alongside SOC 2 Type II for model inference on the Baseten platform. ([source](https://docs.baseten.co/observability/security))

### GDPR / data residency

- RunPod: Stated fully GDPR compliant for data processed in European data center regions, with documented EU procedures for collection, storage, processing and deletion. ([source](https://docs.runpod.io/references/security-and-compliance))
- Baseten: Baseten Cloud supports a GDPR compliance program; residency is set with regional environments or a region on a single deployment. ([source](https://docs.baseten.co/hosting-options/cloud))

### PCI DSS

- RunPod: Claimed for Secure Cloud's vetted infrastructure partners, who meet enterprise standards including SOC 2, ISO 27001 and PCI DSS certifications. Stated of the data-centre partners behind Secure Cloud, not of Runpod Inc. itself. ([source](https://docs.runpod.io/references/security-and-compliance))
- Baseten: PCI DSS - SAQ D listed as a compliance program; a PCI-DSS v4.0.1 AOC for SAQ D Service Provider is published in the Trust Center for request. SAQ D self-assessment AOC, not a QSA Report on Compliance. ([source](https://trust.baseten.co/))

### Encryption at rest

- RunPod: Optional per-Pod volume encryption: the volume disk is encrypted at rest on the host machine, but container disk and network volumes cannot be encrypted. ([source](https://docs.runpod.io/pods/storage/types))
- Baseten: Trust Center control: datastores housing sensitive customer data are encrypted at rest, and transmission over public networks is encrypted. ([source](https://trust.baseten.co/controls))

### Customer-managed keys

- RunPod: Not supported: Runpod stores the volume encryption key, which cannot be retrieved, and the docs state bring your own key is not supported. ([source](https://docs.runpod.io/pods/storage/types))
- Baseten: Runtime OIDC BYOK recipes: a model fetches customer-managed keys at runtime to decrypt encrypted weights and to decrypt requests and encrypt responses. Documented under runtime OIDC use cases, implemented in model code via the truss-examples envelope-encryption recipes. ([source](https://docs.baseten.co/organization/oidc))

### Role-based access control

- RunPod: Four team roles — Basic, Billing, Dev and Admin — each with set permissions; Admin has unrestricted access to members, settings, billing and resources. ([source](https://docs.runpod.io/accounts-billing/manage-accounts))
- Baseten: Role-based access control with three organization roles - Admin, Member and Viewer - in a single-team organization. ([source](https://docs.baseten.co/organization/access))

### Private networking / VPC

- RunPod: Global networking gives each Pod a private IP reachable only by other Pods in the same account; NVIDIA GPU Pods only, in 17 data centers. ([source](https://docs.runpod.io/pods/networking))
- Baseten: Single-tenant Enterprise environments expose endpoints via AWS PrivateLink or Google Cloud Private Service Connect, off the public internet. ([source](https://docs.baseten.co/hosting-options/single-tenant))

### Dedicated infrastructure

- RunPod: Reserved Clusters: dedicated GPU clusters with guaranteed availability, custom configurations and SLA-backed uptime for enterprises scaling to 10,000+ GPUs. ([source](https://www.runpod.io/pricing))
- Baseten: Single-tenant runs workloads in an isolated VPC in Baseten Cloud with compute restricted to your organization; Enterprise, custom pricing. ([source](https://docs.baseten.co/hosting-options/single-tenant))

## Which should you choose?

Pick RunPod if Short experiments and bursty inference — per-second billing and a $0.27 entry card mean an idle hour costs nothing and a small job is genuinely small.

Pick Baseten if Teams that want per-token APIs and dedicated GPUs from one vendor with no plan fee — minute-granularity billing means a model that runs for ten minutes costs ten minutes.

Consider something else: RunPod — The catalogue mixes community and secure capacity, so the cheapest quoted rate is not always the tier a production workload should sit on.

Consider something else: Baseten — Its dedicated H100 works out at $6.50 an hour, roughly three times DeepInfra's, so steady GPU workloads pay a premium for the platform around them.

## Documented by only one

- Free tier: Baseten: No — New workspaces get starting credits for testing; usage beyond them is billed per minute.
- Kind of free offer: Baseten: Free credit
- Storage price: RunPod: Persistent storage from $0.05 per GB per month, in standard and high-performance tiers.
- Input price, top model: Baseten: $1.40 per million input tokens for GLM-5.3, with cached input at $0.14 and output at $4.40. GLM-5.2 Fast, a speed variant, is $2.10.
- Output price, top model: Baseten: $4.40 per million output tokens for GLM-5.3.
- Input price, cheapest model: Baseten: $0.10 per million input tokens for GPT OSS 120B, with output at $0.50. GLM-5.3-Flash is $0.15 in and $0.50 out.
- Context window: Baseten: 1,048K tokens on the DeepSeek V4 line and the GLM 5.2/5.3 family, the widest it serves; 262K on the Kimi K2 models, 200K on GLM 4.7 and 128K on GPT OSS 120B. Baseten's table publishes these in thousands, so the magnitude is 1,048,000 rather than the 1,048,576 a power-of-two reading would give — the vendor's own figure, not a conversion of it.
- Max output tokens: Baseten: 262K tokens on most models, and it is a separate ceiling from the context window rather than the remainder of it: DeepSeek V4 Flash 0731 allows 384K output inside a 1,048K window, while GLM 4.7, Nemotron Ultra and GPT OSS 120B cap output at their full context. The Inkling models are the outlier at 32K.
- Prompt caching: Baseten: Supported
- Rate limits: Baseten: Two limits, requests and tokens per minute, set by account status rather than spend: an unverified Basic account gets 15 RPM and 100,000 TPM, a verified Basic or Pro account 120 RPM and 500,000 TPM, Enterprise custom. Cached input counts toward TPM at full weight even though it is billed cheaper.
- Cheapest GPU, per hour: RunPod: $0.27 per hour for an RTX A5000 with 24GB. RTX 3090 is $0.50, RTX 4090 $0.74, L40S $1.09.
- Billing granularity: RunPod: Per second, with per-hour rates shown for comparison. Clusters can be started with no commitment and scaled to 64 GPUs.
- Reserved pricing: RunPod: Savings plans: a 3-month or 6-month upfront prepaid commitment discounts GPU compute costs; storage stays at standard rates and plans are non-refundable.
- Support on entry paid plan: Baseten: Email and in-app chat support on Basic, which is pay-as-you-go from $0. Pro adds dedicated support on Slack and Zoom.

## Questions this comparison answers

**Entry paid plan: RunPod or Baseten?**

RunPod: No plan tiers: a new account can start with as little as $10 in prepaid credits, and deploying a Pod requires at least one hour of credits for the chosen configuration.
Baseten: $0 per month on the Basic plan — pay as you go, with dedicated deployments, model APIs and training included rather than gated behind a fee.

**Should I choose RunPod or Baseten?**

Pick RunPod if Short experiments and bursty inference — per-second billing and a $0.27 entry card mean an idle hour costs nothing and a small job is genuinely small.
Pick Baseten if Teams that want per-token APIs and dedicated GPUs from one vendor with no plan fee — minute-granularity billing means a model that runs for ten minutes costs ten minutes.
