# Replicate

Canonical: https://inetgeek.com/model-hosting/replicate/

A marketplace of open-source models billed by the second of hardware time, where deploying your own model means paying for the seconds it sits idle too.

## Who it suits

Running someone else's open model without packaging anything. Thousands of public models bill per run — per output image, per thousand tokens, per second of video — so a workload that fires occasionally pays only for the firing, with no instance to keep warm.

Consider something else if You are deploying your own model and it will sit idle. Private models bill the time they spend setting up and waiting as well as working; only fast-booting fine-tunes escape that. The H100 hour is $5.49, against $3.95 on Modal and $3.49 on RunPod.

## iScore

56 / 100 (80% interval 47–65), rank 4 of 4 in model-hosting.

inetGeek's own reading, computed from sourced facts against category peers by disclosed rules. Not a provider claim. Intervals are 80%; a provider with few documented facts is pulled toward the category middle rather than scored on what little is known.
Method: https://inetgeek.com/iscore/. Full output with every criterion's basis: https://inetgeek.com/scores.json. Every model-hosting peer ranked: https://inetgeek.com/model-hosting/replicate/alternatives/

- Price: 37 (15–59), weight 28%, 2 of 5 criteria scored; scored for peers, not here — undocumented or stated without a number or a yes/limited/no: free_offer_type, free_tier, minimum_paid_price
- Capacity and limits: 34 (6–62), weight 22%, 1 of 2 criteria scored; scored for peers, not here — undocumented or stated without a number or a yes/limited/no: regions
- Capabilities: 75 (69–81), weight 22%, 2 of 3 criteria scored; scored for peers, not here — undocumented or stated without a number or a yes/limited/no: prompt_caching
- Trust and compliance: 83 (76–89), weight 17%, 0 of 10 criteria scored; scored for peers, not here — undocumented or stated without a number or a yes/limited/no: customer_managed_keys, dedicated_infrastructure, encryption_at_rest, gdpr_data_residency, hipaa_eligible, iso27001, pci_dss, private_networking, rbac, soc2
- Evidence: 72 (72–72), weight 11%, 4 of 4 criteria scored

## What the documentation says

### Pricing

- Pricing model: Two meters, chosen by the model rather than by you. Public models bill per run — per output image, per thousand tokens, per second of video. Private models you deploy yourself run on dedicated hardware and bill for THE TIME THEY SPEND SETTING UP, THE TIME THEY SPEND IDLE and the time they spend active. Fast-booting fine-tunes are the exception and bill active time only. ([source](https://replicate.com/pricing), read 2026-09-17)

### Infrastructure

- Scales to zero: Not supported — Private models bill idle time; fast-booting fine-tunes do not ([source](https://replicate.com/pricing), read 2026-09-17)
- Autoscaling: Supported — Scales up and down automatically with traffic ([source](https://replicate.com/pricing), read 2026-09-17)

### GPU

- H100 SXM, per GPU-hour: $5.49 per hour for a single Nvidia H100 with 80GB of GPU RAM, published as $0.001525 per second — the rate the per-second meter actually charges. ([source](https://replicate.com/pricing), read 2026-09-17)
- Cheapest GPU, per hour: $0.81 per hour for an Nvidia T4 with 16GB, the cheapest GPU on the hardware list. A CPU-only instance is cheaper still at $0.09 an hour. ([source](https://replicate.com/pricing), read 2026-09-17)
- Largest GPU offered: 160GB on the 2x Nvidia A100 (80GB) instance, the largest GPU RAM figure the hardware table states. Larger multi-GPU A100, H100 and H200 configurations are listed but carry no stated GPU RAM and are sold only with committed spend contracts. ([source](https://replicate.com/pricing), read 2026-09-17)
- Billing granularity: Per second. Every hardware row publishes a per-second rate alongside the hourly one, and the per-second figure is the meter. ([source](https://replicate.com/pricing), read 2026-09-17)

### Support

- Support on entry paid plan: Nothing named on the standard product. A dedicated account manager, priority support, higher GPU limits and performance SLAs are all listed under Enterprise and volume discounts, reached by contacting Replicate rather than by picking a plan. ([source](https://replicate.com/pricing), read 2026-09-17)

## Compared with

- [Replicate vs Modal](https://inetgeek.com/compare/replicate-vs-modal/)
- [Replicate vs RunPod](https://inetgeek.com/compare/replicate-vs-runpod/)
- [Replicate vs Baseten](https://inetgeek.com/compare/replicate-vs-baseten/)
