# DeepInfra

Canonical: https://inetgeek.com/llm-apis/deepinfra/

Serverless inference for open-weight models with published cached-input rates, plus dedicated GPUs billed by the minute.

## Who it suits

Running open-weight models at the lowest published per-token rates here — DeepSeek-V4-Flash at $0.06 per million input is a fraction of what the lab itself charges — with dedicated H100s at $2.20 an hour if throughput demands it.

Consider something else if Only some models carry per-token pricing; the rest bill for inference execution time, which is a different and harder thing to forecast.

## What the documentation says

### Pricing

- Pricing model: Two products on one bill: per-token serverless inference with a separate cached-input rate for each model, and dedicated GPUs billed per minute and invoiced weekly. Models without per-token pricing are billed for inference execution time instead. ([source](https://deepinfra.com/pricing), read 2026-09-06)

### Inference

- Input price, top model: $1.30 per million input tokens for DeepSeek-V4-Pro, the most expensive model in its catalogue. Cached input is $0.10. ([source](https://deepinfra.com/pricing), read 2026-09-06)
- Output price, top model: $2.60 per million output tokens for DeepSeek-V4-Pro. ([source](https://deepinfra.com/pricing), read 2026-09-06)
- Input price, cheapest model: $0.06 per million input tokens for DeepSeek-V4-Flash-0731, with cached input at $0.015 and output at $0.18. DeepSeek itself charges $0.44 at peak for the same checkpoint and Fireworks $0.22. ([source](https://deepinfra.com/pricing), read 2026-09-06)
- Context window: 1024k tokens on the DeepSeek V4 models it serves; 160k on the V3 generation. The window is the model's rather than a platform limit. ([source](https://deepinfra.com/pricing), read 2026-09-06)
- Prompt caching: Supported ([source](https://deepinfra.com/pricing), read 2026-09-06)

### GPU

- H100 SXM, per GPU-hour: $2.20 per GPU-hour for a dedicated H100 80GB — below every dedicated cloud in this dataset. A100 80GB is $0.89, H200 $2.69, B200 $3.69. Billed in minute granularity, invoiced weekly. ([source](https://deepinfra.com/pricing), read 2026-09-06)
- Largest GPU offered: 270GB per GPU on a B300, at $4.89 per GPU-hour. ([source](https://deepinfra.com/pricing), read 2026-09-06)

## Compared with

- [DeepInfra vs Fireworks AI](https://inetgeek.com/compare/deepinfra-vs-fireworks/)
- [DeepInfra vs DeepSeek](https://inetgeek.com/compare/deepinfra-vs-deepseek/)
- [DeepInfra vs Baseten](https://inetgeek.com/compare/deepinfra-vs-baseten/)
- [DeepInfra vs OpenAI](https://inetgeek.com/compare/deepinfra-vs-openai/)
- [DeepInfra vs Novita AI](https://inetgeek.com/compare/deepinfra-vs-novita/)
