4 models across 2 tiers

Gemini API rate card, 2026

Current rates for all 4 Gemini models, including cache read/write, batch discounts, and long-context tiers.

Models

4

Cheapest input

$0.30/MTok

Cheapest output

$2.50/MTok

Cache discount

90% off

Gemini API pricing table

ModelTierInput $/MTokOutput $/MTokCache ReadCache WriteBatchContext
Gemini 3.5 Flash-LiteFast / Budget$0.30$2.50$0.03(90% off)$0.30(no surcharge)$0.15(50% off)1,000K
Gemini 3 FlashFast / Budget$0.50$3.00$0.05(90% off)$0.50(no surcharge)$0.25(50% off)1,000K
Gemini 3.5 FlashFast / Budget$1.50$9.00$0.15(90% off)$1.50(no surcharge)$0.75(50% off)1,000K
Gemini 3.1 ProMid-tier$2.00$12.00$0.20(90% off)$2.00(no surcharge)$1.00(50% off)1,000K

Gemini cache and batch pricing

How caching and batch processing affect your Gemini API costs.

Cache read discount

When a prompt prefix is already cached, the provider charges a reduced rate for those tokens. All Gemini models offer a 90% cache read discount, meaning cached input tokens cost just 10% of the standard input rate.

Cache write cost

Gemini does not charge a cache write surcharge. Writing tokens to cache costs the same as standard input.

Batch discount

All Gemini models offer a 50% batch discount. Batch requests are processed asynchronously and cost 50% of the standard rate.

Gemini long-context pricing

Some Gemini models apply higher rates once the total prompt exceeds a threshold. Both input and output tokens may be re-priced for the entire request.

ModelThresholdInput multiplierOutput multiplier
Gemini 3.1 Pro200.001K tokens2x1.5x

Recent Gemini price changes

Gemini 3.5 Flash input increased

$0.15 to $1.50/MTok

Jul 15, 2026

Gemini 3.5 Flash output increased

$0.60 to $9.00/MTok

Jul 15, 2026

Gemini 3.5 Flash output decreased

$0.75 to $0.60/MTok

Jun 10, 2026

Gemini 3.1 Pro input launched

$2.00/MTok at launch

Apr 22, 2026

View full Gemini pricing history

Track your Gemini spend

See exactly how much you spend on Gemini models and where you can save. Two-line wrapper install, no API keys shared.

Gemini API pricing FAQ

How much does the Gemini API cost?

Gemini API pricing starts at $0.30/MTok for input and $2.50/MTok for output. Prices vary by model tier: flagship models cost more but deliver higher quality, while fast/budget models are significantly cheaper for high-volume workloads.

Does Gemini offer batch discounts?

Yes. Gemini supports batch processing, which sends requests asynchronously at a reduced rate. The batch discount is 50% off standard pricing across all Gemini models.

What is the Gemini cache read discount?

Gemini has a 90% cache read discount. When input tokens are already cached from a previous request, they are charged at 10% of the standard input rate. This matters for applications that send similar prompts repeatedly.

How many Gemini models are available?

There are currently 4 Gemini models available through the API, spanning 2 tiers: Fast / Budget, Mid-tier.

Is there a free tier for the Gemini API?

Gemini API pricing is pay-per-use based on token consumption. There is no permanent free tier, but new accounts typically receive introductory credits. Tokeven helps you track exactly what you spend so nothing is wasted.

Does Gemini charge more for long context?

Yes. Some Gemini models apply a long-context surcharge when the total prompt exceeds a token threshold. Both input and output rates increase for the entire request. Check the long-context pricing table above for exact thresholds and multipliers per model.