OpenAI
Hosted inference for the GPT model family, sold per million tokens with separate short- and long-context rates.
- llm api
- Not documented
- Not documented
- Not documented
Who it suits.
Work that fits inside the short-context rate, where the cheapest tier is $0.20 per million input tokens and the batch tier halves everything above it.
Consider something else if…
Long prompts cross into the long-context rate and roughly double the bill, which is a cliff the flat-rate vendors do not have.
Editorial · Palash Bagchi · approved
What the documentation says.
Each value below links to the page it was read from, with the sentence it came from. Criteria OpenAI does not publish are not listed.
Pricing
Pricing model
Per million tokens, with a second rate above the short-context threshold — long-context input and output cost roughly double. Batch, Flex, Fast and standard tiers are priced separately, and regional processing adds a 10% uplift on models released from March 2026.
Sources (1) →Sources ↓
- Pricing — OpenAI API docs ↗
“Prices per 1M tokens. [...] Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026 [...] Priority processing was renamed Fast mode on July 30, 2026.”
Read 2026-09-06 · official pricing
Inference
Input price, top model
$10 per million input tokens for gpt-6-astra at short context, rising to $20 above the short-context threshold. gpt-5.6-sol is $4 and gpt-5.6-terra $2.
Sources (1) →Sources ↓
- Pricing — OpenAI API docs ↗
“| gpt-6-astra | $10.00 | $1.00 | $12.50 | $50.00 | $20.00 | $2.00 | $25.00 | $75.00 | [...] | gpt-5.6-sol | $4.00 [...] | gpt-5.6-terra | $2.00”
Read 2026-09-06 · official pricing
Output price, top model
$50 per million output tokens for gpt-6-astra at short context, $75 at long context.
Sources (1) →Sources ↓
- Pricing — OpenAI API docs ↗
“| gpt-6-astra | $10.00 | $1.00 | $12.50 | $50.00 | $20.00 | $2.00 | $25.00 | $75.00 |”
Read 2026-09-06 · official pricing
Input price, cheapest model
$0.20 per million input tokens for gpt-5.6-luna at short context, with output at $1.20.
Sources (1) →Sources ↓
- Pricing — OpenAI API docs ↗
“| gpt-5.6-luna | $0.20 | $0.02 | $0.25 | $1.20 | $0.40 | $0.04 | $0.50 | $1.80 |”
Read 2026-09-06 · official pricing
Context window
1,050,000 tokens on gpt-6-astra, of which at most 922,000 may be input.
Sources (1) →Sources ↓
- gpt-6-astra — OpenAI API docs ↗
“- 1,050,000 context window - Maximum input tokens: 922,000”
Read 2026-09-06 · official docs
Max output tokens
128,000 tokens on gpt-6-astra.
Sources (1) →Sources ↓
- gpt-6-astra — OpenAI API docs ↗
“- 128,000 max output tokens”
Read 2026-09-06 · official docs
Prompt caching
Supported
Sources (1) →Sources ↓
- Pricing — OpenAI API docs ↗
“| Model | Short context input | Short context cached input | Short context cache writes | Short context output | [...] | gpt-6-astra | $10.00 | $1.00 | $12.50 | $50.00 |”
Read 2026-09-06 · official pricing
Batch discount
50% off input and output through the Batch API: gpt-6-astra falls from $10/$50 to $5/$25 at short context.
Sources (1) →Sources ↓
- Pricing — OpenAI API docs ↗
“Batch pricing data | Model | Short context input [...] | gpt-6-astra | $5.00 | $0.50 | $6.25 | $25.00 | $10.00 | $1.00 | $12.50 | $37.50 |”
Read 2026-09-06 · official pricing
