Monthly Report6 min read

State of AI API Costs, July 2026

July 15, 2026

Updated 2026-08-16: this report quoted Gemini 3.5 Flash-Lite at $0.25/$1.50 and put Google's cache-read discount at 75%. Neither was ever a real Google price - our registry row was wrong from the day it was written, and Google has published cached reads at 10% of input for as long as we have snapshots. Both are corrected below. The July prices for every other model are left as published.

Updated 2026-08-24: this report listed Claude Sonnet 5 at $3.00/$15.00 with $2/$10 flagged as introductory pricing through Aug 31. Anthropic has cancelled the increase that date pointed at, so $2.00/$10.00 is simply what Claude Sonnet 5 costs - and it is what the model billed at in July as well. The mid-tier line and the sentence after it are corrected below.

July 2026 prices

Three providers have fourteen models with a 200x price gap between cheapest and most expensive. This report breaks down every tier so teams can pick the right model.

Flagship tier: $5-10 per million input tokens

The flagship tier holds steady at $5/MTok input for Claude Opus 4.8 and GPT 5.5. GPT 5.6 Sol matches that $5/MTok price point at launch this quarter.

Claude Fable 5 sits at the premium end: $10/MTok input and $50/MTok output. That makes Fable 5 the most expensive model in the entire registry today.

Output pricing reveals a separate gap worth tracking across providers and workload types. GPT 5.5 and GPT 5.6 Sol charge $30/MTok output, versus Claude Opus 4.8 at $25/MTok. For output-heavy workloads like code generation, that 20% gap compounds over thousands of calls.

Mid tier: the sweet spot

  • -Gemini 3.1 Pro: $2.00/$12.00 (input/output per MTok)
  • -GPT 5.4: $2.50/$15.00
  • -GPT 5.6 Terra: $2.50/$15.00
  • -Claude Sonnet 4.6: $3.00/$15.00
  • -Claude Sonnet 5: $2.00/$10.00

Of the five mid-tier models above, Claude Sonnet 5 and Gemini 3.1 Pro tie for the cheapest input at $2.00 per million tokens. Claude Sonnet 5 takes the output side, at $10.00/MTok against Gemini 3.1 Pro's $12.00/MTok.

Fast tier: where the real savings live

  • -Gemini 3.5 Flash-Lite: $0.30/$2.50
  • -Gemini 3 Flash: $0.50/$3.00
  • -Claude Haiku 4.5: $1.00/$5.00
  • -GPT 5.6 Luna: $1.00/$6.00
  • -Gemini 3.5 Flash: $1.50/$9.00

As of this writing Gemini 3.5 Flash-Lite, at $0.30/MTok input, is the cheapest model across the three providers covered here. For classification and extraction tasks it scores 85/100 confidence at a fraction of flagship cost. (No longer current: GPT 5.6 Luna was cut to $0.20/$1.20 on 2026-08-13, and this report does not cover xAI.)

Cheapest model per task type (min 80 confidence)

  • -Classification: Gemini 3.5 Flash-Lite ($0.30/MTok) or Gemini 3 Flash ($0.50/MTok)
  • -Extraction: Gemini 3.5 Flash-Lite ($0.30/MTok) or Gemini 3 Flash ($0.50/MTok)
  • -Summarization: Gemini 3 Flash ($0.50/MTok)
  • -QA: Gemini 3 Flash ($0.50/MTok)
  • -Code generation: Gemini 3.1 Pro ($2.00/MTok)
  • -Reasoning: Gemini 3.1 Pro ($2.00/MTok)

Key takeaways

  • -Google dominates the budget end of the market with four models under $2/MTok input
  • -Anthropic and OpenAI compete head-to-head in both mid and flagship pricing tiers
  • -The price gap between Flash-Lite ($0.30) and Fable 5 ($10.00) is 33x on input alone
  • -Caching discounts do not vary by provider: Anthropic, Google and OpenAI all price cached reads at 10% of input, i.e. 90% off

Run your numbers through our free teardown at /teardown, or start tracking real spend with a free dashboard.

Track your real costs

Calculators estimate. The dashboard shows what you actually spend and where you can save.