Skip to content
inetGeek

Replicate vs Modal

SEPT 2026 audit

A comparison of Replicate and Modal built from values read directly from each provider's own documentation, with the source recorded against every figure.

Updated 13 criteria comparedSources last checked

The short answer

Choose Replicate if…

Running someone else's open model without packaging anything. Thousands of public models bill per run — per output image, per thousand tokens, per second of video — so a workload that fires occasionally pays only for the firing, with no instance to keep warm.

Editorial · Palash Bagchi · approved

Choose Modal if…

Bursty, self-deployed inference where the bill should follow the work. Compute meters per second and per CPU core — a physical core is $0.0000131 per core-second with a 0.125-core minimum — and the H100 hour is $3.95, the cheapest of the four here that publishes one.

Editorial · Palash Bagchi · approved

Consider something else if…

  • Replicate: You are deploying your own model and it will sit idle. Private models bill the time they spend setting up and waiting as well as working; only fast-booting fine-tunes escape that. The H100 hour is $5.49, against $3.95 on Modal and $3.49 on RunPod.
  • Modal: You want a large GPU named with its memory before committing, or support without an Enterprise contract. The page lists B200 and B300 but states VRAM only for the A100 rows, so the ceiling is unpublished; SOC 2, HIPAA, RBAC and SSO are all Enterprise.

5 sourced criteria separate them — see where, with sources, below.

No affiliate links, sponsored placements or paid rankings appear on this site. Ordering follows the sourced data and the stated criteria.

01.

At a glance.

CriterionReplicateModal
Pricing
Pricing modelTwo meters, chosen by the model rather than by you. Public models bill per run — per output image, per thousand tokens, per second of video. Private models you deploy yourself run on dedicated hardware and bill for THE TIME THEY SPEND SETTING UP, THE TIME THEY SPEND IDLE and the time they spend active. Fast-booting fine-tunes are the exception and bill active time only.A platform fee plus metered compute, and the fee buys a compute allowance rather than access. Starter is $0 a month and includes $30 of compute; Team is $250 a month and includes $100. Compute is billed per second on top of both.
GPU
H100 SXM, per GPU-hour
$5.49 per hour for a single Nvidia H100 with 80GB of GPU RAM, published as $0.001525 per second — the rate the per-second meter actually charges.
$3.95 per hour for an Nvidia H100 SXM5, published as $0.001097 per second. Modal's own worked example on the same page prices a fleet at $3.95 per GPU-hour.
Cheapest GPU, per hour
$0.81 per hour for an Nvidia T4 with 16GB, the cheapest GPU on the hardware list. A CPU-only instance is cheaper still at $0.09 an hour.
$0.59 per hour for an Nvidia T4, published as $0.000164 per second — the cheapest GPU on the resource table.
Billing granularityPer second. Every hardware row publishes a per-second rate alongside the hourly one, and the per-second figure is the meter.Per second, and per CPU core rather than per instance: a physical core is $0.0000131 per core per second with a minimum of 0.125 cores.
Support
Support on entry paid planNothing named on the standard product. A dedicated account manager, priority support, higher GPU limits and performance SLAs are all listed under Enterprise and volume discounts, reached by contacting Replicate rather than by picking a plan.A community Slack on the lower plans; private Slack support and embedded ML engineering services are Enterprise. Support is a plan feature here rather than something every tier carries.

Only criteria both providers publish appear here; a tinted cell marks a real difference. Criteria only one of them documents are listed below, and an absence there means we found no source — not that the feature is missing. How we source this.

The small bar above a value is inetGeek's own lean toward that side — computed from the same facts shown, never a number the provider published. See the picker below "The short answer" to weigh only the criteria you care about.

02.

Where they differ.

Pricing model

Replicate
Two meters, chosen by the model rather than by you. Public models bill per run — per output image, per thousand tokens, per second of video. Private models you deploy yourself run on dedicated hardware and bill for THE TIME THEY SPEND SETTING UP, THE TIME THEY SPEND IDLE and the time they spend active. Fast-booting fine-tunes are the exception and bill active time only.
Modal
A platform fee plus metered compute, and the fee buys a compute allowance rather than access. Starter is $0 a month and includes $30 of compute; Team is $250 a month and includes $100. Compute is billed per second on top of both.
Sources (2) →
  • Pricing – Replicate ↗

    “e they spend setting up; the time they spend idle, waiting for requests; and the time they spend active, processing your requests. If you ge”

    Read 2026-09-17 · official pricing

  • Plans and Pricing — Modal ↗

    “/ mo free PRICING PLANS Starter $0 + compute / month Built for small teams and independent developers looking to level up. Get started with”

    Read 2026-09-17 · official pricing

H100 SXM, per GPU-hour

Replicate
$5.49 per hour for a single Nvidia H100 with 80GB of GPU RAM, published as $0.001525 per second — the rate the per-second meter actually charges.
Modal
$3.95 per hour for an Nvidia H100 SXM5, published as $0.001097 per second. Modal's own worked example on the same page prices a fleet at $3.95 per GPU-hour.
Sources (2) →
  • Pricing – Replicate ↗

    “GPU RAM 160GB RAM 288GB Nvidia H100 GPU gpu-h100 $ 0.001525 /sec $ 5.49 /hr GPU 1x CPU 13x GPU RAM 80GB RAM 144GB Nvidia L40S GPU gpu-l40s”

    Read 2026-09-17 · official pricing

  • Plans and Pricing — Modal ↗

    “H200 SXM $0.001261 / sec Nvidia H100 SXM5 $0.001097 / sec Nvidia RTX PRO 6000 $0.000842 / sec Nvidia A100, 80 GB $0.000694 / sec Nvidia A100”

    Read 2026-09-17 · official pricing

Cheapest GPU, per hour

Replicate
$0.81 per hour for an Nvidia T4 with 16GB, the cheapest GPU on the hardware list. A CPU-only instance is cheaper still at $0.09 an hour.
Modal
$0.59 per hour for an Nvidia T4, published as $0.000164 per second — the cheapest GPU on the resource table.
Sources (2) →
  • Pricing – Replicate ↗

    “x GPU RAM 96GB RAM 144GB Nvidia T4 GPU gpu-t4 $ 0.000225 /sec $ 0.81 /hr GPU 1x CPU 4x GPU RAM 16GB RAM 16GB Additional hardware 4x Nvidia A”

    Read 2026-09-17 · official pricing

  • Plans and Pricing — Modal ↗

    “vidia L4 $0.000222 / sec Nvidia T4 $0.000164 / sec CPU Physical core (2 vCPU equivalent ) $0.0000131 / core / sec *minimum of 0.125 cores pe”

    Read 2026-09-17 · official pricing

Billing granularity

Replicate
Per second. Every hardware row publishes a per-second rate alongside the hourly one, and the per-second figure is the meter.
Modal
Per second, and per CPU core rather than per instance: a physical core is $0.0000131 per core per second with a minimum of 0.125 cores.
Sources (2) →
  • Pricing – Replicate ↗

    “GPU RAM 160GB RAM 288GB Nvidia H100 GPU gpu-h100 $ 0.001525 /sec $ 5.49 /hr GPU 1x CPU 13x GPU RAM 80GB RAM 144GB Nvidia L40S GPU gpu-l40s”

    Read 2026-09-17 · official pricing

  • Plans and Pricing — Modal ↗

    “vidia L4 $0.000222 / sec Nvidia T4 $0.000164 / sec CPU Physical core (2 vCPU equivalent ) $0.0000131 / core / sec *minimum of 0.125 cores pe”

    Read 2026-09-17 · official pricing

Support on entry paid plan

Replicate
Nothing named on the standard product. A dedicated account manager, priority support, higher GPU limits and performance SLAs are all listed under Enterprise and volume discounts, reached by contacting Replicate rather than by picking a plan.
Modal
A community Slack on the lower plans; private Slack support and embedded ML engineering services are Enterprise. Support is a plan feature here rather than something every tier carries.
Sources (2) →
  • Pricing – Replicate ↗

    “uirements, we can offer: Dedicated account manager Priority support Higher GPU limits Performance SLAs Help with onboarding, custom models,”

    Read 2026-09-17 · official pricing

  • Plans and Pricing — Modal ↗

    “nvironment-level budgets Support via private Slack Audit logs, SAML SSO, and HIPAA Credit grants for startups Early-stage startups can get f”

    Read 2026-09-17 · official pricing

When these numbers change, hear about it. Sources are re-checked monthly; a repricing goes out as a short note.

Confirm by email; unsubscribe from any issue. Your address goes to Kit and nowhere else — what we do with it.

03.

Documented by only one.

These criteria are published by one provider and not the other. An absence here means we have not found a source, not that the feature is missing.

Free tierModal: Yes
Kind of free offerModal: Free credit — $30 of compute a month on Starter, recurring rather than one-off
Entry paid planModal: $0 a month on Starter — the plan carries no fee at all, and $30 of monthly compute is included before anything is billed. The first plan with a monthly fee is Team at $250.
Scales to zeroReplicate: Not supported — Private models bill idle time; fast-booting fine-tunes do not
AutoscalingReplicate: Supported — Scales up and down automatically with traffic
Largest GPU offeredReplicate: 160GB on the 2x Nvidia A100 (80GB) instance, the largest GPU RAM figure the hardware table states. Larger multi-GPU A100, H100 and H200 configurations are listed but carry no stated GPU RAM and are sold only with committed spend contracts.
SOC 2Modal: SOC 2 compliance is listed as an Enterprise-plan feature on the pricing page, alongside HIPAA compatibility, audit logs, RBAC and SSO. No report or certificate is linked from the page itself.
HIPAAModal: HIPAA compatibility is listed as an Enterprise-plan feature, grouped with SOC 2 and audit logs. The page says compatibility rather than a signed BAA.
04.

Questions this comparison answers.

Should I choose Replicate or Modal?

Pick Replicate if Running someone else's open model without packaging anything. Thousands of public models bill per run — per output image, per thousand tokens, per second of video — so a workload that fires occasionally pays only for the firing, with no instance to keep warm.

Pick Modal if Bursty, self-deployed inference where the bill should follow the work. Compute meters per second and per CPU core — a physical core is $0.0000131 per core-second with a 0.125-core minimum — and the H100 hour is $3.95, the cheapest of the four here that publishes one.

Provider pages

Related comparisons

Every page here is sourced and dated.