# Novita AI

Canonical: https://inetgeek.com/llm-apis/novita/

Serverless inference across 200+ open models with per-model cache-read rates, plus on-demand and spot GPU instances.

## Who it suits

Breadth — 200+ models under one API with a cache-read rate published next to most of them, and an H100 SXM at $3.39 an hour if the same vendor should run the dedicated side too.

Consider something else if It is rarely the cheapest for any given model: DeepInfra serves DeepSeek-V4-Flash-0731 at $0.06 against Novita's $0.44, and its H100 is $2.20 against Novita's $3.39 on demand.

## What the documentation says

### Pricing

- Pricing model: Per-token serverless inference with a cache-read rate published beside most models, and GPU instances offered at both on-demand and spot rates — the spot column is roughly half. Two models, Ling 3.0 Flash Fin and Sante, are listed at no charge on both input and output. ([source](https://novita.ai/pricing), read 2026-09-07)

### Inference

- Input price, top model: $2.00 per million input tokens for Qwen3.8 Max, the most expensive model in its catalogue, with cache reads at $0.25 and output at $6.00. ([source](https://novita.ai/pricing), read 2026-09-07)
- Output price, top model: $6.00 per million output tokens for Qwen3.8 Max. ([source](https://novita.ai/pricing), read 2026-09-07)
- Input price, cheapest model: $0.075 per million input tokens for GLM 5.3 Flash at short context, rising to $0.15 above the threshold, with cache reads at $0.015 and output at $0.25. Fireworks serves the same model at $0.15. ([source](https://novita.ai/pricing), read 2026-09-07)
- Context window: 1M tokens on the DeepSeek V4 line and GLM 5.3, published as "1M" rather than an exact figure — so it is recorded at 1,000,000 and may in practice be the 1,048,576 other providers state. Qwen3.8 is listed at 977K. ([source](https://novita.ai/pricing), read 2026-09-07)
- Prompt caching: Supported ([source](https://novita.ai/pricing), read 2026-09-07)

### GPU

- H100 SXM, per GPU-hour: $3.39 per GPU-hour for an on-demand H100 SXM 80GB, or $1.70 spot. The page names the card the site's GPU comparison is fixed on, so no substitution is involved. ([source](https://novita.ai/gpus), read 2026-09-07)
- Largest GPU offered: 80GB per GPU on an H100 SXM, the largest card it lists; below it are an L40S 48GB, RTX 5090 32GB and RTX 4090 24GB. ([source](https://novita.ai/gpus), read 2026-09-07)

## Compared with

- [Novita AI vs DeepInfra](https://inetgeek.com/compare/deepinfra-vs-novita/)
- [Novita AI vs Fireworks AI](https://inetgeek.com/compare/fireworks-vs-novita/)
- [Novita AI vs DeepSeek](https://inetgeek.com/compare/deepseek-vs-novita/)
- [Novita AI vs Baseten](https://inetgeek.com/compare/baseten-vs-novita/)
