# Modal vs Baseten

Canonical: https://inetgeek.com/compare/modal-vs-baseten/

Every value below is read from Modal's and Baseten's own documentation. See https://inetgeek.com/methodology/ for how.

## At a glance

| Criterion | Modal | Baseten |
| --- | --- | --- |
| Free tier | Yes | No — New workspaces get starting credits for testing; usage beyond them is billed per minute. |
| Kind of free offer | Free credit — $30 of compute a month on Starter, recurring rather than one-off | Free credit |
| Entry paid plan | $0 a month on Starter — the plan carries no fee at all, and $30 of monthly compute is included before anything is billed. The first plan with a monthly fee is Team at $250. | $0 per month on the Basic plan — pay as you go, with dedicated deployments, model APIs and training included rather than gated behind a fee. |
| Pricing model | A platform fee plus metered compute, and the fee buys a compute allowance rather than access. Starter is $0 a month and includes $30 of compute; Team is $250 a month and includes $100. Compute is billed per second on top of both. | No platform fee: per-token Model APIs for open models, and dedicated deployments billed per minute of compute with volume discounts. The Pro tier buys priority access to high-demand GPUs rather than a lower rate. |
| H100 SXM, per GPU-hour | $3.95 per hour for an Nvidia H100 SXM5, published as $0.001097 per second. Modal's own worked example on the same page prices a fleet at $3.95 per GPU-hour. | $6.50 per hour for a dedicated H100 80GB. Baseten publishes $0.10833 per MINUTE — the hourly figure is that times 60, ours rather than Baseten's, and the page offers an hourly toggle. A100 80GB is $0.06667 a minute and B200 $0.16633. |
| Support on entry paid plan | A community Slack on the lower plans; private Slack support and embedded ML engineering services are Enterprise. Support is a plan feature here rather than something every tier carries. | Email and in-app chat support on Basic, which is pay-as-you-go from $0. Pro adds dedicated support on Slack and Zoom. |
| SOC 2 | SOC 2 compliance is listed as an Enterprise-plan feature on the pricing page, alongside HIPAA compatibility, audit logs, RBAC and SSO. No report or certificate is linked from the page itself. | Maintains SOC 2 Type II certification for the inference platform; policies and certifications are held in the Baseten Trust Center. |
| HIPAA | HIPAA compatibility is listed as an Enterprise-plan feature, grouped with SOC 2 and audit logs. The page says compatibility rather than a signed BAA. | Maintains HIPAA compliance alongside SOC 2 Type II for model inference on the Baseten platform. |

## Where they differ

### Free tier

- Modal: Yes ([source](https://modal.com/pricing))
- Baseten: No — New workspaces get starting credits for testing; usage beyond them is billed per minute. ([source](https://docs.baseten.co/organization/billing))

### Entry paid plan

- Modal: $0 a month on Starter — the plan carries no fee at all, and $30 of monthly compute is included before anything is billed. The first plan with a monthly fee is Team at $250. ([source](https://modal.com/pricing))
- Baseten: $0 per month on the Basic plan — pay as you go, with dedicated deployments, model APIs and training included rather than gated behind a fee. ([source](https://www.baseten.co/pricing/))

### Pricing model

- Modal: A platform fee plus metered compute, and the fee buys a compute allowance rather than access. Starter is $0 a month and includes $30 of compute; Team is $250 a month and includes $100. Compute is billed per second on top of both. ([source](https://modal.com/pricing))
- Baseten: No platform fee: per-token Model APIs for open models, and dedicated deployments billed per minute of compute with volume discounts. The Pro tier buys priority access to high-demand GPUs rather than a lower rate. ([source](https://www.baseten.co/pricing/))

### H100 SXM, per GPU-hour

- Modal: $3.95 per hour for an Nvidia H100 SXM5, published as $0.001097 per second. Modal's own worked example on the same page prices a fleet at $3.95 per GPU-hour. ([source](https://modal.com/pricing))
- Baseten: $6.50 per hour for a dedicated H100 80GB. Baseten publishes $0.10833 per MINUTE — the hourly figure is that times 60, ours rather than Baseten's, and the page offers an hourly toggle. A100 80GB is $0.06667 a minute and B200 $0.16633. ([source](https://www.baseten.co/pricing/))

### Support on entry paid plan

- Modal: A community Slack on the lower plans; private Slack support and embedded ML engineering services are Enterprise. Support is a plan feature here rather than something every tier carries. ([source](https://modal.com/pricing))
- Baseten: Email and in-app chat support on Basic, which is pay-as-you-go from $0. Pro adds dedicated support on Slack and Zoom. ([source](https://www.baseten.co/pricing/))

### SOC 2

- Modal: SOC 2 compliance is listed as an Enterprise-plan feature on the pricing page, alongside HIPAA compatibility, audit logs, RBAC and SSO. No report or certificate is linked from the page itself. ([source](https://modal.com/pricing))
- Baseten: Maintains SOC 2 Type II certification for the inference platform; policies and certifications are held in the Baseten Trust Center. ([source](https://docs.baseten.co/observability/security))

### HIPAA

- Modal: HIPAA compatibility is listed as an Enterprise-plan feature, grouped with SOC 2 and audit logs. The page says compatibility rather than a signed BAA. ([source](https://modal.com/pricing))
- Baseten: Maintains HIPAA compliance alongside SOC 2 Type II for model inference on the Baseten platform. ([source](https://docs.baseten.co/observability/security))

## Which should you choose?

Pick Modal if Bursty, self-deployed inference where the bill should follow the work. Compute meters per second and per CPU core — a physical core is $0.0000131 per core-second with a 0.125-core minimum — and the H100 hour is $3.95, the cheapest of the four here that publishes one.

Pick Baseten if Teams that want per-token APIs and dedicated GPUs from one vendor with no plan fee — minute-granularity billing means a model that runs for ten minutes costs ten minutes.

Consider something else: Modal — You want a large GPU named with its memory before committing, or support without an Enterprise contract. The page lists B200 and B300 but states VRAM only for the A100 rows, so the ceiling is unpublished; SOC 2, HIPAA, RBAC and SSO are all Enterprise.

Consider something else: Baseten — Its dedicated H100 works out at $6.50 an hour, roughly three times DeepInfra's, so steady GPU workloads pay a premium for the platform around them.

## Documented by only one

- Regions: Baseten: Two Baseten regions selectable for regional deployments: us (United States) and eu (European Union); other regions by request to support. GPU model deployments only; Chains, training jobs and shared CPU types cannot select a region.
- Uptime SLA: Baseten: Custom SLAs on Enterprise. No percentage is published on the pricing page.
- Input price, top model: Baseten: $1.40 per million input tokens for GLM-5.3, with cached input at $0.14 and output at $4.40. GLM-5.2 Fast, a speed variant, is $2.10.
- Output price, top model: Baseten: $4.40 per million output tokens for GLM-5.3.
- Input price, cheapest model: Baseten: $0.10 per million input tokens for GPT OSS 120B, with output at $0.50. GLM-5.3-Flash is $0.15 in and $0.50 out.
- Context window: Baseten: 1,048K tokens on the DeepSeek V4 line and the GLM 5.2/5.3 family, the widest it serves; 262K on the Kimi K2 models, 200K on GLM 4.7 and 128K on GPT OSS 120B. Baseten's table publishes these in thousands, so the magnitude is 1,048,000 rather than the 1,048,576 a power-of-two reading would give — the vendor's own figure, not a conversion of it.
- Max output tokens: Baseten: 262K tokens on most models, and it is a separate ceiling from the context window rather than the remainder of it: DeepSeek V4 Flash 0731 allows 384K output inside a 1,048K window, while GLM 4.7, Nemotron Ultra and GPT OSS 120B cap output at their full context. The Inkling models are the outlier at 32K.
- Prompt caching: Baseten: Supported
- Rate limits: Baseten: Two limits, requests and tokens per minute, set by account status rather than spend: an unverified Basic account gets 15 RPM and 100,000 TPM, a verified Basic or Pro account 120 RPM and 500,000 TPM, Enterprise custom. Cached input counts toward TPM at full weight even though it is billed cheaper.
- Cheapest GPU, per hour: Modal: $0.59 per hour for an Nvidia T4, published as $0.000164 per second — the cheapest GPU on the resource table.
- Largest GPU offered: Baseten: 180GB per GPU on a B200, at $0.16633 per minute — $9.98 an hour.
- Billing granularity: Modal: Per second, and per CPU core rather than per instance: a physical core is $0.0000131 per core per second with a minimum of 0.125 cores.
- ISO 27001: Baseten: ISO 27001:2022 listed as a compliance program; an ISO 27001 certificate for Baseten Labs, Inc. is published in the Trust Center for request. Certificate is request-access in the Trust Center, not a public download.
- GDPR / data residency: Baseten: Baseten Cloud supports a GDPR compliance program; residency is set with regional environments or a region on a single deployment.
- PCI DSS: Baseten: PCI DSS - SAQ D listed as a compliance program; a PCI-DSS v4.0.1 AOC for SAQ D Service Provider is published in the Trust Center for request. SAQ D self-assessment AOC, not a QSA Report on Compliance.
- Encryption at rest: Baseten: Trust Center control: datastores housing sensitive customer data are encrypted at rest, and transmission over public networks is encrypted.
- Customer-managed keys: Baseten: Runtime OIDC BYOK recipes: a model fetches customer-managed keys at runtime to decrypt encrypted weights and to decrypt requests and encrypt responses. Documented under runtime OIDC use cases, implemented in model code via the truss-examples envelope-encryption recipes.
- Role-based access control: Baseten: Role-based access control with three organization roles - Admin, Member and Viewer - in a single-team organization.
- Private networking / VPC: Baseten: Single-tenant Enterprise environments expose endpoints via AWS PrivateLink or Google Cloud Private Service Connect, off the public internet.
- Dedicated infrastructure: Baseten: Single-tenant runs workloads in an isolated VPC in Baseten Cloud with compute restricted to your organization; Enterprise, custom pricing.

## Questions this comparison answers

**Entry paid plan: Modal or Baseten?**

Modal: $0 a month on Starter — the plan carries no fee at all, and $30 of monthly compute is included before anything is billed. The first plan with a monthly fee is Team at $250.
Baseten: $0 per month on the Basic plan — pay as you go, with dedicated deployments, model APIs and training included rather than gated behind a fee.

**Free tier: Modal or Baseten?**

Modal: Yes
Baseten: No — New workspaces get starting credits for testing; usage beyond them is billed per minute.

**Should I choose Modal or Baseten?**

Pick Modal if Bursty, self-deployed inference where the bill should follow the work. Compute meters per second and per CPU core — a physical core is $0.0000131 per core-second with a 0.125-core minimum — and the H100 hour is $3.95, the cheapest of the four here that publishes one.
Pick Baseten if Teams that want per-token APIs and dedicated GPUs from one vendor with no plan fee — minute-granularity billing means a model that runs for ten minutes costs ten minutes.
