Cost Savings8 min read

Claude vs GPT vs Gemini: cost per million tokens compared

August 16, 2026

The full pricing table

Here is every Claude, GPT and Gemini model we track, sorted by input price per million tokens. Prices come straight from the model registry, so this table is current rather than a snapshot.

Fast tier ($0.20 to $1.50/MTok input):

  • -GPT 5.6 Luna: $0.20 / $1.20
  • -Gemini 3.5 Flash-Lite: $0.30 / $2.50
  • -Gemini 3 Flash: $0.50 / $3.00
  • -Claude Haiku 4.5: $1.00 / $5.00
  • -Gemini 3.5 Flash: $1.50 / $9.00

Mid tier ($2.00 to $3.00/MTok input):

  • -Claude Sonnet 5: $2.00 / $10.00
  • -GPT 5.6 Terra: $2.00 / $12.00
  • -Gemini 3.1 Pro: $2.00 / $12.00
  • -GPT 5.4: $2.50 / $15.00
  • -Claude Sonnet 4.6: $3.00 / $15.00

Flagship tier ($4.00 to $10.00/MTok input):

  • -GPT 5.6 Sol: $4.00 / $20.00
  • -Claude Opus 4.8: $5.00 / $25.00
  • -Claude Opus 5: $5.00 / $25.00
  • -GPT 5.5: $5.00 / $30.00
  • -Claude Fable 5.1: $10.00 / $50.00
  • -Claude Fable 5: $10.00 / $50.00

xAI is the fourth provider we track, and its Grok family is deliberately not in the tables above - this comparison is scoped to the three providers in the title. Grok prices are on the model pages and in the teardown; we publish no quality scores for them yet, because our weekly benchmark source does not cover Grok.

Cheapest provider by task type

The cheapest model per task type depends on the minimum quality threshold you need. At 80+ confidence score:

  • -Classification: GPT 5.6 Luna at $0.20/MTok
  • -Extraction: GPT 5.6 Luna at $0.20/MTok
  • -Summarization: GPT 5.6 Luna at $0.20/MTok
  • -Q&A: GPT 5.6 Luna at $0.20/MTok
  • -Code generation: Claude Sonnet 5 at $2.00/MTok (tied on price with GPT 5.6 Terra and Gemini 3.1 Pro; Claude Sonnet 5 scores highest)
  • -Code review: Claude Sonnet 5 at $2.00/MTok (tied on price with GPT 5.6 Terra and Gemini 3.1 Pro; Claude Sonnet 5 scores highest)
  • -Reasoning: Claude Sonnet 5 at $2.00/MTok (tied on price with Gemini 3.1 Pro and GPT 5.6 Terra; Claude Sonnet 5 scores highest)
  • -Creative writing: Claude Sonnet 5 at $2.00/MTok (tied on price with Gemini 3.1 Pro and GPT 5.6 Terra; Claude Sonnet 5 scores highest)

The budget end is not one provider's territory any more. GPT 5.6 Luna at $0.20/MTok now undercuts every Gemini model after OpenAI's August cuts, and it clears the 80-confidence bar on all four high-volume tasks. In the mid tier the three providers land on the same $2.00/MTok list price, so the choice there is a quality call, not a price one.

Cache discounts no longer separate providers

OpenAI and Google price cached input at 0.10x list - 90% off - on every model, and Anthropic does on all but two (Claude Fable 5.1 discounts further still, at 98%). What still differs is the cache-WRITE fee: 25% on Anthropic, 25% on the GPT 5.6 family, nothing on GPT 5.5, GPT 5.4 or any Gemini model.

At an 80% hit rate, with the misses paying the write fee, a million prompt tokens costs:

  • -Claude Opus 4.8: $1.65 (from $5.00 list)
  • -GPT 5.5: $1.40 (from $5.00 list)
  • -Gemini 3.1 Pro: $0.56 (from $2.00 list)

Because the discount is identical, caching does not reorder models that differ on list price. It does split models that tie on it: Claude Opus 4.8 and GPT 5.5 both list at $5.00, and once you are paying to write cache, GPT 5.5 is the cheaper prompt. Pick on list price and output price, then check the write fee if your prefix churns.

Output pricing matters more than you think

For output-heavy workloads (code generation, creative writing, blog writing), output tokens often outnumber input tokens 2-4x. Output pricing determines total cost:

  • -A 2,000-input / 8,000-output call on Claude Opus 4.8: $0.21/call
  • -Same call on GPT 5.5: $0.25/call (19% more, on $30.00 output)
  • -Same call on Claude Sonnet 5: $0.084/call

For code generation, Claude Sonnet 5 is the best value in the mid tier, and no longer a temporary one: its $2.00 / $10.00 launch rate became the standard price when Anthropic cancelled the increase that had been scheduled for September 2026.

The bottom line

No single provider is cheapest across all workloads. The optimal strategy is routing: send classification to the cheapest model that clears your quality bar, code review to a mid-tier model, and only genuinely hard reasoning to a flagship. A team that routes by task type instead of defaulting to one model saves 40-60% on their total AI bill.

Compare all models for your specific workload at /teardown, or start tracking your multi-provider spend with the free dashboard.

Track your real costs

Calculators estimate. The dashboard shows what you actually spend and where you can save.