Monthly Report6 min read

State of AI API Costs, July 2026

July 15, 2026

The pricing landscape in July 2026

Three providers offer fourteen models with a 200x price gap between cheapest and most expensive. This report breaks down every tier so teams can pick the right model.

Flagship tier: $5-10 per million input tokens

The flagship tier holds steady at $5/MTok input for Claude Opus 4.8 and GPT 5.5. GPT 5.6 Sol matches that $5/MTok price point at launch this quarter.

Claude Fable 5 sits at the premium end: $10/MTok input and $50/MTok output. That makes Fable 5 the most expensive model in the entire registry today.

Output pricing reveals a separate gap worth tracking across providers and workload types. GPT 5.5 and GPT 5.6 Sol charge $30/MTok output, versus Claude Opus at $25/MTok. For output-heavy workloads like code generation, that 20% gap compounds over thousands of calls.

Mid tier: the sweet spot

  • -Gemini 3.1 Pro: $2.00/$12.00 (input/output per MTok)
  • -GPT 5.4: $2.50/$15.00
  • -GPT 5.6 Terra: $2.50/$15.00
  • -Claude Sonnet 4.6: $3.00/$15.00
  • -Claude Sonnet 5: $3.00/$15.00 (introductory pricing of $2/$10 through Aug 31)

Gemini 3.1 Pro is the cheapest mid-tier option at $2.00 per million input tokens. Claude Sonnet 5 introductory pricing makes it temporarily competitive at $2/$10 through August 2026.

Fast tier: where the real savings live

  • -Gemini 3.5 Flash-Lite: $0.25/$1.50
  • -Gemini 3 Flash: $0.50/$3.00
  • -Claude Haiku 4.5: $1.00/$5.00
  • -GPT 5.6 Luna: $1.00/$6.00
  • -Gemini 3.5 Flash: $1.50/$9.00

Gemini 3.5 Flash-Lite at $0.25/MTok input is the cheapest model across all three providers. For classification and extraction tasks it scores 85/100 confidence at a fraction of flagship cost.

Cheapest model per task type (min 80 confidence)

  • -Classification: Gemini 3.5 Flash-Lite ($0.25/MTok) or Gemini 3 Flash ($0.50/MTok)
  • -Extraction: Gemini 3.5 Flash-Lite ($0.25/MTok) or Gemini 3 Flash ($0.50/MTok)
  • -Summarization: Gemini 3 Flash ($0.50/MTok)
  • -QA: Gemini 3 Flash ($0.50/MTok)
  • -Code generation: Gemini 3.1 Pro ($2.00/MTok)
  • -Reasoning: Gemini 3.1 Pro ($2.00/MTok)

Key takeaways

  • -Google dominates the budget end of the market with four models under $2/MTok input
  • -Anthropic and OpenAI compete head-to-head in both mid and flagship pricing tiers
  • -The price gap between Flash-Lite ($0.25) and Fable 5 ($10.00) is 40x on input alone
  • -Caching discounts vary by provider: Anthropic offers 90%, Google 75%, and OpenAI 50-90% off reads

Run your numbers through our free teardown at /teardown, or start tracking real spend with a free dashboard.

Track your real costs

Calculators estimate. The dashboard shows what you actually spend and where you can save.