Cost Savings8 min read

Claude vs GPT vs Gemini: cost per million tokens compared

July 22, 2026

The full pricing table (July 2026)

Here is every model from all three providers, sorted by input price per million tokens.

Fast tier ($0.25 to $1.50/MTok input):

  • -Gemini 3.5 Flash-Lite: $0.25 input / $1.50 output
  • -Gemini 3 Flash: $0.50 / $3.00
  • -Claude Haiku 4-5: $1.00 / $5.00
  • -GPT 5.6 Luna: $1.00 / $6.00
  • -Gemini 3.5 Flash: $1.50 / $9.00

Mid tier ($2.00 to $3.00/MTok input):

  • -Gemini 3.1 Pro: $2.00 / $12.00
  • -Claude Sonnet 5 (intro pricing): $2.00 / $10.00
  • -GPT 5.4: $2.50 / $15.00
  • -GPT 5.6 Terra: $2.50 / $15.00
  • -Claude Sonnet 4-6: $3.00 / $15.00

Flagship tier ($5.00 to $10.00/MTok input):

  • -Claude Opus 4-8: $5.00 / $25.00
  • -GPT 5.5: $5.00 / $30.00
  • -GPT 5.6 Sol: $5.00 / $30.00
  • -Claude Fable 5: $10.00 / $50.00

Cheapest provider by task type

The cheapest model per task type depends on the minimum quality threshold you need. At 80+ confidence score:

  • -Classification: Gemini 3.5 Flash-Lite at $0.25/MTok
  • -Extraction: Gemini 3.5 Flash-Lite at $0.25/MTok
  • -Summarization: Gemini 3 Flash at $0.50/MTok
  • -Q&A: Gemini 3 Flash at $0.50/MTok
  • -Code generation: Gemini 3.1 Pro at $2.00/MTok
  • -Code review: Gemini 3.1 Pro at $2.00/MTok
  • -Reasoning: GPT 5.4 at $2.50/MTok
  • -Creative writing: GPT 5.4 at $2.50/MTok

Google dominates the budget end of the market. For quality-sensitive workloads, Anthropic and OpenAI compete head-to-head in mid and flagship tiers.

Cache discounts change the math

Anthropic and OpenAI offer 90% off cached input tokens. Google offers 75% off. This matters for workloads with repeated system prompts.

At 80% cache hit rate on a 5,000-token system prompt:

  • -Anthropic effective input: $0.50/MTok (for a $5.00 model)
  • -OpenAI effective input: $0.50/MTok (for a $5.00 model)
  • -Google effective input: $0.50/MTok (for a $2.00 model)

With high cache rates, the gap between providers narrows significantly. Anthropic's deeper discount offsets its higher list price on input-heavy, cache-friendly workloads.

Output pricing matters more than you think

For output-heavy workloads (code generation, creative writing, blog writing), output tokens often outnumber input tokens 2-4x. Output pricing determines total cost:

  • -A 2,000-input / 8,000-output call on Opus: $0.21/call
  • -Same call on GPT 5.5: $0.25/call (20% more expensive due to $30 output)
  • -Same call on Sonnet 5 (intro): $0.084/call (60% cheaper than Opus)

For code generation, Claude Sonnet 5 at introductory pricing is the best value in mid-tier until August 31.

The bottom line

No single provider is cheapest across all workloads. The optimal strategy is routing: send classification to Flash-Lite, code review to Sonnet, and complex reasoning to Opus. A team that routes by task type instead of defaulting to one model saves 40-60% on their total AI bill.

Compare all models for your specific workload at /teardown, or start tracking your multi-provider spend with the free dashboard.

Track your real costs

Calculators estimate. The dashboard shows what you actually spend and where you can save.