Novita AI
Serverless inference across 200+ open models with per-model cache-read rates, plus on-demand and spot GPU instances.
- LLM APIs
- Not documented
- Not documented
- Not documented
Who it suits.
Breadth — 200+ models under one API with a cache-read rate published next to most of them, and an H100 SXM at $3.39 an hour if the same vendor should run the dedicated side too.
Consider something else if…
It is rarely the cheapest for any given model: DeepInfra serves DeepSeek-V4-Flash-0731 at $0.06 against Novita's $0.44, and its H100 is $2.20 against Novita's $3.39 on demand.
Editorial · Palash Bagchi · approved
What the documentation says.
Each value below links to the page it was read from, with the sentence it came from. Criteria Novita AI does not publish are not listed.
Pricing
Pricing model
Per-token serverless inference with a cache-read rate published beside most models, and GPU instances offered at both on-demand and spot rates — the spot column is roughly half. Two models, Ling 3.0 Flash Fin and Sante, are listed at no charge on both input and output.
Sources (1) →Sources ↓
- Pricing — Novita AI ↗
“Ling 3.0 Flash Fin 256K Free Free [...] Ling 3.0 Flash Sante 256K Free Free [...] On-Demand SPOT”
Read 2026-09-07 · official pricing
Inference
Input price, top model
$2.00 per million input tokens for Qwen3.8 Max, the most expensive model in its catalogue, with cache reads at $0.25 and output at $6.00.
Sources (1) →Sources ↓
- Pricing — Novita AI ↗
“Qwen3.8 Max 977K $2 /Mt · Cache Read $0.25 /Mt $6 /Mt”
Read 2026-09-07 · official pricing
Output price, top model
$6.00 per million output tokens for Qwen3.8 Max.
Sources (1) →Sources ↓
- Pricing — Novita AI ↗
“Qwen3.8 Max 977K $2 /Mt · Cache Read $0.25 /Mt $6 /Mt”
Read 2026-09-07 · official pricing
Input price, cheapest model
$0.075 per million input tokens for GLM 5.3 Flash at short context, rising to $0.15 above the threshold, with cache reads at $0.015 and output at $0.25. Fireworks serves the same model at $0.15.
Sources (1) →Sources ↓
- Pricing — Novita AI ↗
“GLM 5.3 Flash 1M $0.075 /Mt · Cache Read $0.015 /Mt $0.15 /Mt · Cache Read $0.03 /Mt $0.25 /Mt $0.5 /Mt”
Read 2026-09-07 · official pricing
Context window
1M tokens on the DeepSeek V4 line and GLM 5.3, published as "1M" rather than an exact figure — so it is recorded at 1,000,000 and may in practice be the 1,048,576 other providers state. Qwen3.8 is listed at 977K.
Sources (1) →Sources ↓
- Pricing — Novita AI ↗
“DeepSeek V4 Pro 0813 1M [...] GLM 5.3 1M [...] Qwen3.8 Max 977K”
Read 2026-09-07 · official pricing
Prompt caching
Supported
Sources (1) →Sources ↓
- Pricing — Novita AI ↗
“DeepSeek V4 Pro 0813 1M $1.32 /Mt · Cache Read $0.132 /Mt $3.96 /Mt [...] GLM 4.7 Flash 195K $0.07 /Mt · Cache Read $0.01 /Mt $0.4 /Mt”
Read 2026-09-07 · official pricing
GPU
H100 SXM, per GPU-hour
$3.39 per GPU-hour for an on-demand H100 SXM 80GB, or $1.70 spot. The page names the card the site's GPU comparison is fixed on, so no substitution is involved.
Sources (1) →Sources ↓
- GPU Instance — Novita AI ↗
“On-Demand SPOT H100 SXM 80GB 80 GB VRAM 3.39/hr/GPU 1.70/hr/GPU”
Read 2026-09-07 · official pricing
Largest GPU offered
80GB per GPU on an H100 SXM, the largest card it lists; below it are an L40S 48GB, RTX 5090 32GB and RTX 4090 24GB.
Sources (1) →Sources ↓
- GPU Instance — Novita AI ↗
“H100 SXM 80GB 80 GB VRAM 3.39/hr/GPU 1.70/hr/GPU [...] NVIDIA L40S 48GB 48 GB VRAM 0.55/hr/GPU”
Read 2026-09-07 · official pricing
When these numbers change, hear about it. Sources are re-checked monthly; a repricing goes out as a short note.
Novita AI in the news
- Tencent HY3 Tops OpenRouter Charts: Base URL Swap Runs It Free in Codex CLI - Tech Times
- Novita AI Launches Sandbox To Secure Autonomous Agent Systems At Enterprise Scale - Pulse 2.0
- Novita AI Ranked as the Best Performing & Reliable Inference Layer - PR Newswire
- Fastino Launches Pioneer, the First Agent for Fine-tuning and Inference of LLMs - Yahoo Finance
