Skip to content
inetGeek
replicate.com · Model hosting and inference
SEPT 2026 audit

A marketplace of open-source models billed by the second of hardware time, where deploying your own model means paying for the seconds it sits idle too.

Visit Replicate documentation
Compare Replicate with

No affiliate links, sponsored placements or paid rankings appear on this site. Ordering follows the sourced data and the stated criteria.

CategoryModel hosting and inference

Updated 8 criteria comparedSources last checked

01.Evaluation guide

Who it suits

Running someone else's open model without packaging anything. Thousands of public models bill per run — per output image, per thousand tokens, per second of video — so a workload that fires occasionally pays only for the firing, with no instance to keep warm.

Consider something else if…

You are deploying your own model and it will sit idle. Private models bill the time they spend setting up and waiting as well as working; only fast-booting fine-tunes escape that. The H100 hour is $5.49, against $3.95 on Modal and $3.49 on RunPod.

Editorial · Palash Bagchi · approved

Each value below links to the page it was read from, with the sentence it came from. Criteria Replicate does not publish are not listed.

Replicate pricing

1 criteria
Pricing modelTwo meters, chosen by the model rather than by you. Public models bill per run — per output image, per thousand tokens, per second of video. Private models you deploy yourself run on dedicated hardware and bill for THE TIME THEY SPEND SETTING UP, THE TIME THEY SPEND IDLE and the time they spend active. Fast-booting fine-tunes are the exception and bill active time only.
Sources (1) →
  • Pricing – Replicate ↗

    “e they spend setting up; the time they spend idle, waiting for requests; and the time they spend active, processing your requests. If you ge”

    Read 2026-09-17 · official pricing

Infrastructure

2 criteria
Scales to zeroNot supported — Private models bill idle time; fast-booting fine-tunes do not
Sources (1) →
  • Pricing – Replicate ↗

    “e they spend setting up; the time they spend idle, waiting for requests; and the time they spend active, processing your requests. If you ge”

    Read 2026-09-17 · official pricing

AutoscalingSupported — Scales up and down automatically with traffic
Sources (1) →
  • Pricing – Replicate ↗

    “ou get a ton of traffic, we automatically scale up and down to handle the demand. For fast booting fine-tunes you'll only be billed for the”

    Read 2026-09-17 · official pricing

GPU

4 criteria
H100 SXM, per GPU-hour$5.49 per hour for a single Nvidia H100 with 80GB of GPU RAM, published as $0.001525 per second — the rate the per-second meter actually charges.
Sources (1) →
  • Pricing – Replicate ↗

    “GPU RAM 160GB RAM 288GB Nvidia H100 GPU gpu-h100 $ 0.001525 /sec $ 5.49 /hr GPU 1x CPU 13x GPU RAM 80GB RAM 144GB Nvidia L40S GPU gpu-l40s”

    Read 2026-09-17 · official pricing

Cheapest GPU, per hour$0.81 per hour for an Nvidia T4 with 16GB, the cheapest GPU on the hardware list. A CPU-only instance is cheaper still at $0.09 an hour.
Sources (1) →
  • Pricing – Replicate ↗

    “x GPU RAM 96GB RAM 144GB Nvidia T4 GPU gpu-t4 $ 0.000225 /sec $ 0.81 /hr GPU 1x CPU 4x GPU RAM 16GB RAM 16GB Additional hardware 4x Nvidia A”

    Read 2026-09-17 · official pricing

Largest GPU offered160GB on the 2x Nvidia A100 (80GB) instance, the largest GPU RAM figure the hardware table states. Larger multi-GPU A100, H100 and H200 configurations are listed but carry no stated GPU RAM and are sold only with committed spend contracts.
Sources (1) →
  • Pricing – Replicate ↗

    “x GPU RAM 80GB RAM 144GB 2x Nvidia A100 (80GB) GPU gpu-a100-large-2x $ 0.002800 /sec $ 10.08 /hr GPU 2x CPU 20x GPU RAM 160GB RAM 288GB Nvid”

    Read 2026-09-17 · official pricing

Billing granularityPer second. Every hardware row publishes a per-second rate alongside the hourly one, and the per-second figure is the meter.
Sources (1) →
  • Pricing – Replicate ↗

    “GPU RAM 160GB RAM 288GB Nvidia H100 GPU gpu-h100 $ 0.001525 /sec $ 5.49 /hr GPU 1x CPU 13x GPU RAM 80GB RAM 144GB Nvidia L40S GPU gpu-l40s”

    Read 2026-09-17 · official pricing

Support

1 criteria
Support on entry paid planNothing named on the standard product. A dedicated account manager, priority support, higher GPU limits and performance SLAs are all listed under Enterprise and volume discounts, reached by contacting Replicate rather than by picking a plan.
Sources (1) →
  • Pricing – Replicate ↗

    “uirements, we can offer: Dedicated account manager Priority support Higher GPU limits Performance SLAs Help with onboarding, custom models,”

    Read 2026-09-17 · official pricing

When these numbers change, hear about it. Sources are re-checked monthly; a repricing goes out as a short note.

Confirm by email; unsubscribe from any issue. Your address goes to Kit and nowhere else — what we do with it.

Compare Replicate withAll alternatives, ranked →

Every page here is sourced and dated.