GPU compute
Each provider is profiled from its own documentation, with the source recorded against every figure. The unit here is fixed by the hardware rather than the vendor: an H100 SXM 80GB is the same card everywhere, so the hourly rate for one is directly comparable — and every price names the card it was read for.
| Provider | H100 SXM, per GPU-hour | Largest card offered (VRAM) |
|---|---|---|
| Baseten | $6.50 per hour for a dedicated H100 80GB. Baseten publishes $0.10833 per MINUTE — the hourly figure is that times 60, ours rather than Baseten's, and the page offers an hourly toggle. A100 80GB is $0.06667 a minute and B200 $0.16633. | 180GB per GPU on a B200, at $0.16633 per minute — $9.98 an hour. |
| CoreWeave | $6.16 per GPU-hour for H100 SXM 80GB. CoreWeave sells the 8-GPU HGX H100 node at $49.24/hour on demand — the per-GPU figure is that divided by eight, ours rather than CoreWeave's, and you cannot rent one card. | 288GB per GPU on B300-class hardware; the largest listed on-demand node is HGX B200 at 180GB per GPU, eight GPUs, $68.80/hour. |
| DeepInfra | $2.20 per GPU-hour for a dedicated H100 80GB — below every dedicated cloud in this dataset. A100 80GB is $0.89, H200 $2.69, B200 $3.69. Billed in minute granularity, invoiced weekly. | 270GB per GPU on a B300, at $4.89 per GPU-hour. |
| Fireworks AI | $7.00 per hour for an H100 80GB on demand, or $8.00 on the shorter commitment shown beside it. The same rate applies to an H200 141GB. | 180GB per GPU on a B200, at $10.00 an hour on demand. |
| Hyperstack | $3.20 per hour for H100 SXM 80GB on demand. The NVLink variant is $2.60 and the PCIe H100 is $2.50. | 288GB per GPU on B300 at $7.40 an hour; B200 at 192GB is $6.00. |
| Lambda | $4.29 per GPU-hour for H100 SXM 80GB on a single-GPU instance, falling to $3.99 on an 8-GPU node. The PCIe variant is $3.29. | 180GB per GPU on B200 SXM6, from $6.69 per GPU-hour on an 8-GPU node to $6.99 on a single GPU. |
| Novita AI | $3.39 per GPU-hour for an on-demand H100 SXM 80GB, or $1.70 spot. The page names the card the site's GPU comparison is fixed on, so no substitution is involved. | 80GB per GPU on an H100 SXM, the largest card it lists; below it are an L40S 48GB, RTX 5090 32GB and RTX 4090 24GB. |
| RunPod | $3.49 per hour for H100 SXM 80GB, single card. The PCIe variant is $2.89 and H100 NVL is $3.19. | 288GB per GPU on B300; B200 at 180GB is $6.79 an hour and H200 at 141GB is $4.59. |
| Vast.ai | Median $2.16 per hour for H100 SXM 80GB, with a floor of $1.33 from the cheapest host. Vast is a marketplace — the floor is one host's offer at one moment and the median is what the market is actually charging, so both move without anyone announcing anything. This figure is a snapshot, not a rate card. | 192GB per GPU on B200, from $4.38 an hour with a median of $7.50. |
Which to consider
Teams that want per-token APIs and dedicated GPUs from one vendor with no plan fee — minute-granularity billing means a model that runs for ten minutes costs ten minutes.
Training runs that want a whole HGX node and move large datasets — free egress and free I/O operations remove the charge that usually dominates a data-heavy job.
Running open-weight models at the lowest published per-token rates here — DeepSeek-V4-Flash at $0.06 per million input is a fraction of what the lab itself charges — with dedicated H100s at $2.20 an hour if throughput demands it.
Running open-weight models without operating GPUs — the same DeepSeek and Qwen checkpoints the labs publish, served per token, with a path to dedicated hardware if throughput demands it.
Steady European workloads that can commit — reserved H200 at $2.79 against $3.99 on demand is the clearest published discount in this set, and egress is free.
Renting exactly the instance shape you need — Lambda publishes 1x, 2x, 4x and 8x plans for the same card, so a single H100 is a real product rather than a fraction of a node.
Breadth — 200+ models under one API with a cache-read rate published next to most of them, and an H100 SXM at $3.39 an hour if the same vendor should run the dedicated side too.
Short experiments and bursty inference — per-second billing and a $0.27 entry card mean an idle hour costs nothing and a small job is genuinely small.
Price-sensitive training and batch inference that can tolerate a marketplace — the median H100 rate is well under half what the managed clouds charge, and interruptible capacity goes lower still.
LLM APIs
GPU compute
GPU compute in the news
- Jensen Huang Calls Nvidia Chips a ‘Revenue-Generating Asset’ as Older H100 Rentals Reach $3.28 an Hour - tipranks.com
- Bitget launches H100 and B200 GPU compute perpetuals - Invezz
- Bitget Lists Nvidia H100 And B200 Perpetuals For Retail Crypto Traders - Yellow.com
- Google Puts $12/Hr Price on Ironwood TPU vs Nvidia [2026] - shattered.io
