Skip to content
inetGeek

Google Gemini API

Hosted inference for the Gemini model family, with a free tier and per-million-token paid pricing.

Visit Google Gemini API documentation
Compare Google Gemini API with
ai.google.dev

No affiliate links, sponsored placements or paid rankings appear on this site. Ordering follows the sourced data and the stated criteria.

Updated 10 criteria compared

Category
llm api
Free tier
Yes
Entry paid plan
Not documented
Regions
Not documented
01.

Who it suits.

Getting to a working prototype without a card, then staying on the cheapest frontier-adjacent rates in this set — $0.75 per million input tokens is well under half what Anthropic and OpenAI charge for their top models.

Consider something else if…

The rates are dated: they roughly double on 1 January 2027, free-tier content is used to improve Google's products, and rate limits are not published anywhere you can read before signing up.

Editorial · Palash Bagchi · approved

02.

What the documentation says.

Each value below links to the page it was read from, with the sentence it came from. Criteria Google Gemini API does not publish are not listed.

Pricing

Free tier

Yes

Sources (1) →
  • Gemini Developer API Pricing ↗

    “Free input & output tokens [...] Limited access to certain models [...] Google AI Studio access [...] Content used to improve our products [...] Get started for Free”

    Read 2026-09-06 · official pricing

Pricing model

Per million tokens, with dated price changes published in advance — the current Gemini 3.x rates are explicitly promotional through 31 December 2026 and roughly double on 1 January 2027. A free tier exists, and its content is used to improve Google's products where the paid tier's is not.

Sources (1) →
  • Gemini Developer API Pricing ↗

    “Paid Tier, per 1M tokens in USD [...] $0.75 through December 31, 2026. $1.50 starting January 1, 2027. [...] Used to improve our products | Yes | No”

    Read 2026-09-06 · official pricing

Inference

Input price, top model

$0.75 per million input tokens for Gemini 3.8 Flash through 31 December 2026, then $1.50. Gemini 3.1 Pro, still preview, is $2.00 for prompts up to 200k tokens and $4.00 above.

Sources (1) →
  • Gemini Developer API Pricing ↗

    “Gemini 3.8 Flash [...] Input price | Free of charge | $0.75 through December 31, 2026. $1.50 starting January 1, 2027. [...] Gemini 3.1 Pro Preview [...] Input price | Not available | $2.00, prompts <= 200k tokens $4.00, prompts > 200k tokens”

    Read 2026-09-06 · official pricing

Output price, top model

$3.75 per million output tokens for Gemini 3.8 Flash through 31 December 2026, then $7.50. Thinking tokens are billed as output.

Sources (1) →
  • Gemini Developer API Pricing ↗

    “Output price (including thinking tokens) | Free of charge | $3.75 through December 31, 2026. $7.50 starting January 1, 2027.”

    Read 2026-09-06 · official pricing

Input price, cheapest model

$0.30 per million input tokens for Gemini 3.5 Flash-Lite, with output at $2.50.

Sources (1) →
  • Gemini Developer API Pricing ↗

    “Gemini 3.5 Flash-Lite [...] Input price | Free of charge | $0.30 (text / image / video / audio) Output price (including thinking tokens) | Free of charge | $2.50”

    Read 2026-09-06 · official pricing

Context window

1,048,576 tokens for Gemini 3.8 Flash. Read from the Vertex AI model card: the Gemini API docs build their model cards client-side and publish no figure a fetch can read.

Sources (1) →

Max output tokens

65,536 tokens for Gemini 3.8 Flash, and Google counts thinking tokens as output.

Sources (1) →

Prompt caching

Supported

Sources (1) →
  • Gemini Developer API Pricing ↗

    “Context caching price | Free of charge | $0.075 through December 31, 2026. $0.15 starting January 1, 2027. $0.50 / 1,000,000 tokens per hour (storage price)”

    Read 2026-09-06 · official pricing

Batch discount

50% off through the Batch API, which is a paid-tier feature.

Sources (1) →

Rate limits

Tied to a usage tier — Free, then Tier 1/2/3 unlocked by cumulative spend — and enforced as requests per minute, input tokens per minute and requests per day. Google publishes no numbers in the docs: limits are viewed in AI Studio and are stated not to be guaranteed.

Sources (1) →
  • Rate limits — Gemini API ↗

    “Rate limits depend on a variety of factors (such as your usage tier) and can be viewed in Google AI Studio. [...] Specified rate limits are not guaranteed and actual capacity may vary. [...] Usage tier | Qualification | Billing tier cap Free | Active project or free trial | N/A Tier 1 | Set up and link an active billing account | $250”

    Read 2026-09-06 · official docs