Skip to content
inetGeek

Novita AI

Serverless inference across 200+ open models with per-model cache-read rates, plus on-demand and spot GPU instances.

Visit Novita AI documentation
Compare Novita AI with
novita.ai

No affiliate links, sponsored placements or paid rankings appear on this site. Ordering follows the sourced data and the stated criteria.

Updated 8 criteria comparedSources last checked

Category
LLM APIs
Free tier
Not documented
Entry paid plan
Not documented
Regions
Not documented
01.

Who it suits.

Breadth — 200+ models under one API with a cache-read rate published next to most of them, and an H100 SXM at $3.39 an hour if the same vendor should run the dedicated side too.

Consider something else if…

It is rarely the cheapest for any given model: DeepInfra serves DeepSeek-V4-Flash-0731 at $0.06 against Novita's $0.44, and its H100 is $2.20 against Novita's $3.39 on demand.

Editorial · Palash Bagchi · approved

02.

What the documentation says.

Each value below links to the page it was read from, with the sentence it came from. Criteria Novita AI does not publish are not listed.

Pricing

Pricing model

Per-token serverless inference with a cache-read rate published beside most models, and GPU instances offered at both on-demand and spot rates — the spot column is roughly half. Two models, Ling 3.0 Flash Fin and Sante, are listed at no charge on both input and output.

Sources (1) →
  • Pricing — Novita AI ↗

    “Ling 3.0 Flash Fin 256K Free Free [...] Ling 3.0 Flash Sante 256K Free Free [...] On-Demand SPOT”

    Read 2026-09-07 · official pricing

Inference

Input price, top model

$2.00 per million input tokens for Qwen3.8 Max, the most expensive model in its catalogue, with cache reads at $0.25 and output at $6.00.

Sources (1) →

Output price, top model

$6.00 per million output tokens for Qwen3.8 Max.

Sources (1) →

Input price, cheapest model

$0.075 per million input tokens for GLM 5.3 Flash at short context, rising to $0.15 above the threshold, with cache reads at $0.015 and output at $0.25. Fireworks serves the same model at $0.15.

Sources (1) →
  • Pricing — Novita AI ↗

    “GLM 5.3 Flash 1M $0.075 /Mt · Cache Read $0.015 /Mt $0.15 /Mt · Cache Read $0.03 /Mt $0.25 /Mt $0.5 /Mt”

    Read 2026-09-07 · official pricing

Context window

1M tokens on the DeepSeek V4 line and GLM 5.3, published as "1M" rather than an exact figure — so it is recorded at 1,000,000 and may in practice be the 1,048,576 other providers state. Qwen3.8 is listed at 977K.

Sources (1) →
  • Pricing — Novita AI ↗

    “DeepSeek V4 Pro 0813 1M [...] GLM 5.3 1M [...] Qwen3.8 Max 977K”

    Read 2026-09-07 · official pricing

Prompt caching

Supported

Sources (1) →
  • Pricing — Novita AI ↗

    “DeepSeek V4 Pro 0813 1M $1.32 /Mt · Cache Read $0.132 /Mt $3.96 /Mt [...] GLM 4.7 Flash 195K $0.07 /Mt · Cache Read $0.01 /Mt $0.4 /Mt”

    Read 2026-09-07 · official pricing

GPU

H100 SXM, per GPU-hour

$3.39 per GPU-hour for an on-demand H100 SXM 80GB, or $1.70 spot. The page names the card the site's GPU comparison is fixed on, so no substitution is involved.

Sources (1) →

Largest GPU offered

80GB per GPU on an H100 SXM, the largest card it lists; below it are an L40S 48GB, RTX 5090 32GB and RTX 4090 24GB.

Sources (1) →
  • GPU Instance — Novita AI ↗

    “H100 SXM 80GB 80 GB VRAM 3.39/hr/GPU 1.70/hr/GPU [...] NVIDIA L40S 48GB 48 GB VRAM 0.55/hr/GPU”

    Read 2026-09-07 · official pricing

When these numbers change, hear about it. Sources are re-checked monthly; a repricing goes out as a short note.

Confirm by email; unsubscribe from any issue. Your address goes to Kit and nowhere else — what we do with it.

Novita AI in the news

Headlines from other publications, not checked by us · collected

Selected by a search on this company, refreshed weekly. Nobody here has verified these claims — where something material is confirmed, it is written into the page above with its source, the way every other fact on this site is.

Every page here is sourced and dated.