DeepInfra vs Novita AI
A comparison of DeepInfra and Novita AI built from values read directly from each provider's own documentation, with the source recorded against every figure.
The short answer
- Pick DeepInfra if
- Running open-weight models at the lowest published per-token rates here — DeepSeek-V4-Flash at $0.06 per million input is a fraction of what the lab itself charges — with dedicated H100s at $2.20 an hour if throughput demands it.
- Pick Novita AI if
- Breadth — 200+ models under one API with a cache-read rate published next to most of them, and an H100 SXM at $3.39 an hour if the same vendor should run the dedicated side too.
7 sourced criteria separate them:
- Pricing model
- DeepInfra: Two products on one bill: per-token serverless inference with a separate cached-input rate for each model, and dedicated GPUs billed per minute and invoiced weekly. Models without per-token pricing are billed for inference execution time instead.
- Novita AI: Per-token serverless inference with a cache-read rate published beside most models, and GPU instances offered at both on-demand and spot rates — the spot column is roughly half. Two models, Ling 3.0 Flash Fin and Sante, are listed at no charge on both input and output.
- Input price, top model
- DeepInfra: $1.30 per million input tokens for DeepSeek-V4-Pro, the most expensive model in its catalogue. Cached input is $0.10.
- Novita AI: $2.00 per million input tokens for Qwen3.8 Max, the most expensive model in its catalogue, with cache reads at $0.25 and output at $6.00.
- Output price, top model
- DeepInfra: $2.60 per million output tokens for DeepSeek-V4-Pro.
- Novita AI: $6.00 per million output tokens for Qwen3.8 Max.
- Input price, cheapest model
- DeepInfra: $0.06 per million input tokens for DeepSeek-V4-Flash-0731, with cached input at $0.015 and output at $0.18. DeepSeek itself charges $0.44 at peak for the same checkpoint and Fireworks $0.22.
- Novita AI: $0.075 per million input tokens for GLM 5.3 Flash at short context, rising to $0.15 above the threshold, with cache reads at $0.015 and output at $0.25. Fireworks serves the same model at $0.15.
- Context window
- DeepInfra: 1024k tokens on the DeepSeek V4 models it serves; 160k on the V3 generation. The window is the model's rather than a platform limit.
- Novita AI: 1M tokens on the DeepSeek V4 line and GLM 5.3, published as "1M" rather than an exact figure — so it is recorded at 1,000,000 and may in practice be the 1,048,576 other providers state. Qwen3.8 is listed at 977K.
- H100 SXM, per GPU-hour
- DeepInfra: $2.20 per GPU-hour for a dedicated H100 80GB — below every dedicated cloud in this dataset. A100 80GB is $0.89, H200 $2.69, B200 $3.69. Billed in minute granularity, invoiced weekly.
- Novita AI: $3.39 per GPU-hour for an on-demand H100 SXM 80GB, or $1.70 spot. The page names the card the site's GPU comparison is fixed on, so no substitution is involved.
- Largest GPU offered
- DeepInfra: 270GB per GPU on a B300, at $4.89 per GPU-hour.
- Novita AI: 80GB per GPU on an H100 SXM, the largest card it lists; below it are an L40S 48GB, RTX 5090 32GB and RTX 4090 24GB.
Pick the criteria you care about. The chart counts how many of them lean toward each provider — the same read as scanning the bars below, just totalled for the ones you chose.
DeepInfra 0
Novita AI 0
At a glance.
| Criterion | DeepInfra | Novita AI |
|---|---|---|
| Pricing | ||
| Pricing model | Two products on one bill: per-token serverless inference with a separate cached-input rate for each model, and dedicated GPUs billed per minute and invoiced weekly. Models without per-token pricing are billed for inference execution time instead. | Per-token serverless inference with a cache-read rate published beside most models, and GPU instances offered at both on-demand and spot rates — the spot column is roughly half. Two models, Ling 3.0 Flash Fin and Sante, are listed at no charge on both input and output. |
| Inference | ||
| Input price, top model | $1.30 per million input tokens for DeepSeek-V4-Pro, the most expensive model in its catalogue. Cached input is $0.10. | $2.00 per million input tokens for Qwen3.8 Max, the most expensive model in its catalogue, with cache reads at $0.25 and output at $6.00. |
| Output price, top model | $2.60 per million output tokens for DeepSeek-V4-Pro. | $6.00 per million output tokens for Qwen3.8 Max. |
| Input price, cheapest model | $0.06 per million input tokens for DeepSeek-V4-Flash-0731, with cached input at $0.015 and output at $0.18. DeepSeek itself charges $0.44 at peak for the same checkpoint and Fireworks $0.22. | $0.075 per million input tokens for GLM 5.3 Flash at short context, rising to $0.15 above the threshold, with cache reads at $0.015 and output at $0.25. Fireworks serves the same model at $0.15. |
| Context window | 1024k tokens on the DeepSeek V4 models it serves; 160k on the V3 generation. The window is the model's rather than a platform limit. | 1M tokens on the DeepSeek V4 line and GLM 5.3, published as "1M" rather than an exact figure — so it is recorded at 1,000,000 and may in practice be the 1,048,576 other providers state. Qwen3.8 is listed at 977K. |
| Prompt caching | Supported | Supported |
| GPU | ||
| H100 SXM, per GPU-hour | $2.20 per GPU-hour for a dedicated H100 80GB — below every dedicated cloud in this dataset. A100 80GB is $0.89, H200 $2.69, B200 $3.69. Billed in minute granularity, invoiced weekly. | $3.39 per GPU-hour for an on-demand H100 SXM 80GB, or $1.70 spot. The page names the card the site's GPU comparison is fixed on, so no substitution is involved. |
| Largest GPU offered | 270GB per GPU on a B300, at $4.89 per GPU-hour. | 80GB per GPU on an H100 SXM, the largest card it lists; below it are an L40S 48GB, RTX 5090 32GB and RTX 4090 24GB. |
Where they differ.
Pricing model
- DeepInfra
- Two products on one bill: per-token serverless inference with a separate cached-input rate for each model, and dedicated GPUs billed per minute and invoiced weekly. Models without per-token pricing are billed for inference execution time instead.
- Novita AI
- Per-token serverless inference with a cache-read rate published beside most models, and GPU instances offered at both on-demand and spot rates — the spot column is roughly half. Two models, Ling 3.0 Flash Fin and Sante, are listed at no charge on both input and output.
Sources (2) →Sources ↓
- Pricing — DeepInfra ↗
“Some of our langauge models offer per token pricing. Most other models are billed for inference execution time. [...] Billed in minute granularity Invoiced weekly”
Read 2026-09-06 · official pricing
- Pricing — Novita AI ↗
“Ling 3.0 Flash Fin 256K Free Free [...] Ling 3.0 Flash Sante 256K Free Free [...] On-Demand SPOT”
Read 2026-09-07 · official pricing
Input price, top model
- DeepInfra
- $1.30 per million input tokens for DeepSeek-V4-Pro, the most expensive model in its catalogue. Cached input is $0.10.
- Novita AI
- $2.00 per million input tokens for Qwen3.8 Max, the most expensive model in its catalogue, with cache reads at $0.25 and output at $6.00.
Sources (2) →Sources ↓
- Pricing — DeepInfra ↗
“DeepSeek-V4-Pro 1024k $1.30 / $0.10 cached $2.60”
Read 2026-09-06 · official pricing
- Pricing — Novita AI ↗
“Qwen3.8 Max 977K $2 /Mt · Cache Read $0.25 /Mt $6 /Mt”
Read 2026-09-07 · official pricing
Output price, top model
- DeepInfra
- $2.60 per million output tokens for DeepSeek-V4-Pro.
- Novita AI
- $6.00 per million output tokens for Qwen3.8 Max.
Sources (2) →Sources ↓
- Pricing — DeepInfra ↗
“DeepSeek-V4-Pro 1024k $1.30 / $0.10 cached $2.60”
Read 2026-09-06 · official pricing
- Pricing — Novita AI ↗
“Qwen3.8 Max 977K $2 /Mt · Cache Read $0.25 /Mt $6 /Mt”
Read 2026-09-07 · official pricing
Input price, cheapest model
- DeepInfra
- $0.06 per million input tokens for DeepSeek-V4-Flash-0731, with cached input at $0.015 and output at $0.18. DeepSeek itself charges $0.44 at peak for the same checkpoint and Fireworks $0.22.
- Novita AI
- $0.075 per million input tokens for GLM 5.3 Flash at short context, rising to $0.15 above the threshold, with cache reads at $0.015 and output at $0.25. Fireworks serves the same model at $0.15.
Sources (2) →Sources ↓
- Pricing — DeepInfra ↗
“DeepSeek-V4-Flash-0731 1024k $0.06 / $0.015 cached $0.18”
Read 2026-09-06 · official pricing
- Pricing — Novita AI ↗
“GLM 5.3 Flash 1M $0.075 /Mt · Cache Read $0.015 /Mt $0.15 /Mt · Cache Read $0.03 /Mt $0.25 /Mt $0.5 /Mt”
Read 2026-09-07 · official pricing
Context window
1024k tokens on the DeepSeek V4 models it serves; 160k on the V3 generation. The window is the model's rather than a platform limit.
1M tokens on the DeepSeek V4 line and GLM 5.3, published as "1M" rather than an exact figure — so it is recorded at 1,000,000 and may in practice be the 1,048,576 other providers state. Qwen3.8 is listed at 977K.
Sources (2) →Sources ↓
- Pricing — DeepInfra ↗
“DeepSeek-V4-Pro 1024k [...] DeepSeek-V3.2 160k $0.26 / $0.13 cached $0.38”
Read 2026-09-06 · official pricing
- Pricing — Novita AI ↗
“DeepSeek V4 Pro 0813 1M [...] GLM 5.3 1M [...] Qwen3.8 Max 977K”
Read 2026-09-07 · official pricing
H100 SXM, per GPU-hour
- DeepInfra
- $2.20 per GPU-hour for a dedicated H100 80GB — below every dedicated cloud in this dataset. A100 80GB is $0.89, H200 $2.69, B200 $3.69. Billed in minute granularity, invoiced weekly.
- Novita AI
- $3.39 per GPU-hour for an on-demand H100 SXM 80GB, or $1.70 spot. The page names the card the site's GPU comparison is fixed on, so no substitution is involved.
Sources (2) →Sources ↓
- Pricing — DeepInfra ↗
“A100 80GB $0.89 / GPU-hour H100 80GB $2.20 / GPU-hour H200 141GB $2.69 / GPU-hour B200 180GB $3.69 / GPU-hour B300 270GB $4.89 / GPU-hour [...] Billed in minute granularity Invoiced weekly”
Read 2026-09-06 · official pricing
- GPU Instance — Novita AI ↗
“On-Demand SPOT H100 SXM 80GB 80 GB VRAM 3.39/hr/GPU 1.70/hr/GPU”
Read 2026-09-07 · official pricing
Largest GPU offered
270GB per GPU on a B300, at $4.89 per GPU-hour.
80GB per GPU on an H100 SXM, the largest card it lists; below it are an L40S 48GB, RTX 5090 32GB and RTX 4090 24GB.
Sources (2) →Sources ↓
- Pricing — DeepInfra ↗
“B300 270GB $4.89 / GPU-hour”
Read 2026-09-06 · official pricing
- GPU Instance — Novita AI ↗
“H100 SXM 80GB 80 GB VRAM 3.39/hr/GPU 1.70/hr/GPU [...] NVIDIA L40S 48GB 48 GB VRAM 0.55/hr/GPU”
Read 2026-09-07 · official pricing
When these numbers change, hear about it. Sources are re-checked monthly; a repricing goes out as a short note.
Which should you choose?
Choose DeepInfra if…
Running open-weight models at the lowest published per-token rates here — DeepSeek-V4-Flash at $0.06 per million input is a fraction of what the lab itself charges — with dedicated H100s at $2.20 an hour if throughput demands it.
Editorial · Palash Bagchi · approved
Choose Novita AI if…
Breadth — 200+ models under one API with a cache-read rate published next to most of them, and an H100 SXM at $3.39 an hour if the same vendor should run the dedicated side too.
Editorial · Palash Bagchi · approved
Consider something else if…
- DeepInfra: Only some models carry per-token pricing; the rest bill for inference execution time, which is a different and harder thing to forecast.
- Novita AI: It is rarely the cheapest for any given model: DeepInfra serves DeepSeek-V4-Flash-0731 at $0.06 against Novita's $0.44, and its H100 is $2.20 against Novita's $3.39 on demand.
Questions this comparison answers.
Should I choose DeepInfra or Novita AI?
Pick DeepInfra if Running open-weight models at the lowest published per-token rates here — DeepSeek-V4-Flash at $0.06 per million input is a fraction of what the lab itself charges — with dedicated H100s at $2.20 an hour if throughput demands it.
Pick Novita AI if Breadth — 200+ models under one API with a cache-read rate published next to most of them, and an H100 SXM at $3.39 an hour if the same vendor should run the dedicated side too.
