Skip to content
inetGeek

Anthropic

Hosted inference for the Claude model family, sold per million tokens with prompt caching and a batch tier.

Visit Anthropic documentation
Compare Anthropic with
www.anthropic.com

No affiliate links, sponsored placements or paid rankings appear on this site. Ordering follows the sourced data and the stated criteria.

Updated 8 criteria compared

Category
llm api
Free tier
Not documented
Entry paid plan
Not documented
Regions
Not documented
01.

Who it suits.

Teams that want the full 1M-token window at the standard per-token rate rather than a long-context surcharge, and that can use caching and batch to cut a frontier-model bill roughly in half.

Consider something else if…

Its cheapest model is $1 per million input tokens, several times what the budget tiers at Google, OpenAI and DeepSeek charge, so high-volume simple work is priced against you.

Editorial · Palash Bagchi · approved

02.

What the documentation says.

Each value below links to the page it was read from, with the sentence it came from. Criteria Anthropic does not publish are not listed.

Pricing

Pricing model

Per million tokens, priced separately for input and output and per model, with multipliers stacked on top: prompt caching, a 50% batch discount, a 1.1x uplift for US-only inference on Claude 4.6 and later, and a fast-mode premium.

Sources (1) →
  • Pricing — Claude Docs ↗

    “This page provides detailed pricing information for Anthropic's models and features. All prices are in USD. [...] These multipliers stack with other pricing modifiers, including the Batch API discount and data residency. [...] specifying US-only inference through the inference_geo parameter incurs a 1.1x multiplier on all token pricing categories”

    Read 2026-09-06 · official pricing

Inference

Input price, top model

$10 per million input tokens for Claude Fable 5.1, the most capable model. Claude Opus 5 is $5 and Claude Sonnet 5 is $2.

Sources (1) →
  • Pricing — Claude Docs ↗

    “Claude Fable 5.1 | $10 / MTok | $12.50 / MTok | $20 / MTok | $0.25 / MTok | $50 / MTok [...] Claude Opus 5 | $5 / MTok [...] Claude Sonnet 5 | $2 / MTok”

    Read 2026-09-06 · official pricing

Output price, top model

$50 per million output tokens for Claude Fable 5.1. Claude Opus 5 is $25 and Claude Sonnet 5 is $10.

Sources (1) →
  • Pricing — Claude Docs ↗

    “Claude Fable 5.1 | $10 / MTok | [...] | $50 / MTok [...] Claude Opus 5 | [...] $25 / MTok [...] Claude Sonnet 5 | [...] $10 / MTok”

    Read 2026-09-06 · official pricing

Input price, cheapest model

$1 per million input tokens for Claude Haiku 4.5, the cheapest current model, with output at $5.

Sources (1) →
  • Pricing — Claude Docs ↗

    “Claude Haiku 4.5 | $1 / MTok | $1.25 / MTok | $2 / MTok | $0.10 / MTok | $5 / MTok”

    Read 2026-09-06 · official pricing

Context window

1M tokens on Claude Fable 5.1, Claude Opus 5 and Claude Sonnet 5; 200K on Claude Haiku 4.5. Anthropic states the full 1M window is billed at standard per-token rates.

Sources (1) →

Max output tokens

128K tokens on the top three models, 64K on Claude Haiku 4.5. The Batch API supports up to 300K output tokens on several models behind a beta header.

Sources (1) →
  • Models overview — Claude Docs ↗

    “Max output | 128K tokens | 128K tokens | 128K tokens | 64K tokens [...] On the Message Batches API, Claude Opus 5, Claude Sonnet 5 [...] support up to 300k output tokens with the output-300k-2026-03-24 beta header.”

    Read 2026-09-06 · official docs

Prompt caching

Supported

Sources (1) →
  • Pricing — Claude Docs ↗

    “Prompt caching reduces costs and latency by reusing previously processed portions of your prompt across API calls. [...] 5-minute cache write | 1.25x base input price [...] 1-hour cache write | 2x base input price [...] A cache hit costs 10% of the standard input price”

    Read 2026-09-06 · official pricing

Batch discount

50% off both input and output tokens through the Batch API, for asynchronous processing.

Sources (1) →
  • Pricing — Claude Docs ↗

    “The Batch API allows asynchronous processing of large volumes of requests with a 50% discount on both input and output tokens.”

    Read 2026-09-06 · official pricing