Gemini/Fast / Budget

Gemini 3.1 Flash-Lite

Gemini 3.1 Flash-Lite is Gemini's fastest and most affordable model, priced at $0.25/MTok input and $1.50/MTok output with a 1,048.576K context window. Backfilled on 2026-09-23. Prices confirmed by the provider pricing page and the LiteLLM cross-check. Release date not published in a machine-readable source. Not yet benchmarked, so never recommended.

Input $/MTok

$0.25

Output $/MTok

$1.50

Context window

1,048.576K

Cache read discount

90% off

Gemini 3.1 Flash-Lite pricing breakdown

RateInput $/MTokOutput $/MTok
Standard$0.25$1.50
Cached read (90% off input)$0.03-
Cached write (no surcharge)$0.25-
Batch (50% off)$0.13$0.75

All rates in USD per million tokens. Prices verified against provider documentation.

Gemini 3.1 Flash-Lite task benchmarks

Confidence scores (0-100) and cross-provider rank for each task type. Rank 1 is best across all 72 models.

Not benchmarked. The third-party benchmark source used for confidence scores does not cover Gemini models yet. Scores will appear here once independent benchmarks are published.

Gemini 3.1 Flash-Lite cost by use case

Estimated cost per API call using default token counts for each workload.

Use caseInput tokensOutput tokensCost / call
Code Generation2,0004,000$0.006500
Unit Test Generation3,0005,000$0.008250
SQL Generation1,5001,000$0.001875
API Generation3,0006,000$0.009750
Code Refactoring4,0004,000$0.007000
Code Review5,0002,000$0.004250
PR Review8,0003,000$0.006500
Security Review6,0002,500$0.005250
Bug Detection5,0002,000$0.004250
Document Summarization10,0001,000$0.004000
PDF Summarization15,0001,500$0.006000
Meeting Notes8,0002,000$0.005000

Monthly cost projections

Projected monthly spend at different call volumes, using code generation as a representative workload (2,000 input / 4,000 output tokens per call).

Calls / monthWithout cachingWith caching (90% read discount)
1K$6.5$6.19
10K$65$61.85
100K$650$618.5
500K$3,250$3,092.5
1M$6,500$6,185

Caching estimate assumes 70% of input tokens are cache reads. Your actual cache hit rate will depend on prompt structure and reuse.

Fast / Budget alternatives to Gemini 3.1 Flash-Lite

Same-tier models from other providers, sorted by input price.

ModelProviderInput $/MTokOutput $/MTokInput savings
GPT 5 NanoGPT$0.05$0.4080% cheaper
GPT 4.1 NanoGPT$0.10$0.4060% cheaper
GPT 6 LunaGPT$0.10$0.5060% cheaper
GPT 4o MiniGPT$0.15$0.6040% cheaper
GPT 5.6 LunaGPT$0.20$1.2020% cheaper
GPT 5.4 NanoGPT$0.20$1.2520% cheaper
GPT 5 MiniGPT$0.25$2.00Same price
GPT 4.1 MiniGPT$0.40$1.6060% more
GPT 3.5 TurboGPT$0.50$1.50100% more
GPT 3.5 Turbo (0125)GPT$0.50$1.50100% more
GPT 5.4 MiniGPT$0.75$4.50200% more
Claude Haiku 4.5Claude$1.00$5.00300% more
Grok Build 0.1Grok$1.00$2.00300% more
o3 MiniGPT$1.10$4.40340% more
o4 MiniGPT$1.10$4.40340% more

Track your Gemini 3.1 Flash-Lite spend

See exactly how much you spend on Gemini 3.1 Flash-Lite and where you can save. Two-line wrapper install, no API keys shared.