DeepSeek
Hosted inference for the DeepSeek model family, priced per million tokens with peak and off-peak rates.
- llm api
- Not documented
- Not documented
- Not documented
Who it suits.
Batch-shaped work that can run outside weekday mornings UTC, where off-peak halves an already low rate and a cache hit costs about 3% of a cache miss.
Consider something else if…
Capacity is expressed as concurrent requests rather than tokens per minute, which is a different thing to plan against, and the peak window lands on European working hours.
Editorial · Palash Bagchi · approved
What the documentation says.
Each value below links to the page it was read from, with the sentence it came from. Criteria DeepSeek does not publish are not listed.
Pricing
Pricing model
Per million tokens, with every rate doubling at peak: off-peak is half price, and peak is 01:00–04:00 and 06:00–10:00 UTC on weekdays. Cache hits are priced roughly 30x below cache misses.
Sources (1) →Sources ↓
- Models & Pricing — DeepSeek API Docs ↗
“The prices listed below are in units of per 1M tokens. [...] Off-peak rates are half of the peak rates. Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday (all other hours are off-peak).”
Read 2026-09-06 · official pricing
Inference
Input price, top model
$1.32 per million input tokens for deepseek-v4-pro at peak on a cache miss, halving to $0.66 off-peak. A cache hit is $0.044 peak.
Sources (1) →Sources ↓
- Models & Pricing — DeepSeek API Docs ↗
“1M INPUT TOKENS (CACHE MISS) OFF-PEAK $0.22 $0.66 $0.22 PEAK $0.44 $1.32 $0.44 [...] 1M INPUT TOKENS (CACHE HIT) OFF-PEAK $0.007 $0.022 $0.007 PEAK $0.014 $0.044 $0.014”
Read 2026-09-06 · official pricing
Output price, top model
$3.96 per million output tokens for deepseek-v4-pro at peak, $1.98 off-peak.
Sources (1) →Sources ↓
- Models & Pricing — DeepSeek API Docs ↗
“1M OUTPUT TOKENS OFF-PEAK $0.66 $1.98 $0.66 PEAK $1.32 $3.96 $1.32”
Read 2026-09-06 · official pricing
Input price, cheapest model
$0.44 per million input tokens for deepseek-v4-flash at peak on a cache miss, $0.22 off-peak, with output at $1.32 peak.
Sources (1) →Sources ↓
- Models & Pricing — DeepSeek API Docs ↗
“MODEL deepseek-v4-flash [...] 1M INPUT TOKENS (CACHE MISS) OFF-PEAK $0.22 [...] PEAK $0.44 [...] 1M OUTPUT TOKENS [...] PEAK $1.32”
Read 2026-09-06 · official pricing
Context window
1M tokens on every model in the current line.
Sources (1) →Sources ↓
- Models & Pricing — DeepSeek API Docs ↗
“CONTEXT LENGTH 1M”
Read 2026-09-06 · official pricing
Max output tokens
384K tokens maximum.
Sources (1) →Sources ↓
- Models & Pricing — DeepSeek API Docs ↗
“MAX OUTPUT MAXIMUM: 384K”
Read 2026-09-06 · official pricing
Prompt caching
Supported
Sources (1) →Sources ↓
- Models & Pricing — DeepSeek API Docs ↗
“1M INPUT TOKENS (CACHE HIT) OFF-PEAK $0.007 $0.022 $0.007 PEAK $0.014 $0.044 $0.014”
Read 2026-09-06 · official pricing
Rate limits
Concurrency rather than requests per minute: 500 concurrent requests on deepseek-v4-pro and 2,500 on deepseek-v4-flash, per account rather than per key. Exceeding it returns HTTP 429, and expansion is free on request.
Sources (1) →Sources ↓
- Rate Limit & Isolation — DeepSeek API Docs ↗
“Concurrency Limit | 500 | 2500 | 2500 [...] Concurrency limits are calculated at the account level, regardless of which API Key is used [...] when the concurrency limit is exceeded, you will receive an HTTP 429 error code [...] There is no additional cost for capacity expansion.”
Read 2026-09-06 · official docs
