4 models across 2 tiers
Google has models in the mid and fast tiers: mid-tier (Gemini 3.1 Pro) and fast (Gemini 3.5 Flash, 3 Flash, 3.5 Flash-Lite). All four Gemini models list under $2.00/MTok input. GPT 5.6 Luna took the outright budget floor in August 2026. Google has a 90% cache read discount with no cache write surcharge, which makes it the cheapest family to maintain a prefix that changes often.
Models
4
Cheapest input
$0.30/MTok
Cache discount
90% off
Max context
1,000K
| Model | Tier | Input $/MTok | Output $/MTok | Context | Best task | Released |
|---|---|---|---|---|---|---|
| Gemini 3.5 Flash-Lite | Fast / Budget | $0.30 | $2.50 | 1,000K | Classification (85/100) | Jun 2026 |
| Gemini 3 Flash | Fast / Budget | $0.50 | $3.00 | 1,000K | Classification (88/100) | Mar 2026 |
| Gemini 3.5 Flash | Fast / Budget | $1.50 | $9.00 | 1,000K | Classification (92/100) | May 2026 |
| Gemini 3.1 Pro | Mid-tier | $2.00 | $12.00 | 1,000K | Reasoning (90/100) | Apr 2026 |
Estimated cost per API call across common workloads, using default token counts for each use case.
Confidence scores (0-100) across all task types. Higher is better.
Confidence scores are a third-party benchmark snapshot from week 29 of 2026. This snapshot is 8 weeks old, so treat the scores as directional and validate against your own evals.
| Task | Flash-Lite | Flash | Flash | Pro |
|---|---|---|---|---|
| Code Generation | 55/100 | 65/100 | 68/100 | 86/100 |
| Code Review | 52/100 | 62/100 | 65/100 | 84/100 |
| Summarization | 79/100 | 83/100 | 85/100 | 87/100 |
| Q&A | 78/100 | 82/100 | 84/100 | 87/100 |
| Extraction | 83/100 | 86/100 | 90/100 | 85/100 |
| Reasoning | 50/100 | 58/100 | 60/100 | 90/100 |
| Classification | 85/100 | 88/100 | 92/100 | 86/100 |
| Creative Writing | 45/100 | 52/100 | 55/100 | 84/100 |
Gemini 3.5 Flash input increased
$0.15 to $1.50/MTok
Gemini 3.5 Flash output increased
$0.60 to $9.00/MTok
Gemini 3.5 Flash output decreased
$0.75 to $0.60/MTok
Gemini 3.1 Pro input launched
$2.00/MTok at launch
See exactly how much you spend on Gemini models and where you can save. Two-line wrapper install, no API keys shared.