Editorial5 min read

Are AI Tokens Getting Cheaper? The Data Says Yes

July 1, 2026

The changelog tells the story

We track every pricing change across all three major providers in a single registry. The data from the first half of 2026 reveals a clear and consistent downward trend.

Price cuts (6 events)

  • -Apr 22: Gemini 3.1 Pro launched at $2.00/MTok input (new model, competitive price)
  • -May 15: GPT 5.5 input dropped from $7.50 to $6.00 (20% cut)
  • -Jun 9: Claude Fable 5 launched at $10.00/MTok (new flagship, premium tier)
  • -Jun 10: Gemini 3.5 Flash output dropped from $0.75 to $0.60 (20% cut)
  • -Jun 30: Claude Sonnet 5 launched at $3.00/MTok (new model, mid-tier pricing)
  • -Jul 9: GPT 5.6 Sol launched at $5.00/MTok (new flagship, matched Opus pricing)

Price increases (1 event)

  • -May 28: Claude Opus 4.8 input increased from $4.50 to $5.00 (11% increase)

The pattern

Six downward moves versus one upward move tell a clear story about market direction. New models launch at price points that undercut or match their predecessors across all three providers.

The one increase, Opus moving from $4.50 to $5.00, likely reflects Anthropic repositioning it as premium. That repositioning happened just ahead of the Sonnet 5 launch at the mid-tier price point.

What is driving prices down

  • -Competition: three providers compete aggressively on price in the mid and fast tiers each quarter
  • -Hardware efficiency: inference costs continue to fall as providers optimize their serving infrastructure at scale
  • -Market segmentation: providers launch tiered model families like Sol/Terra/Luna and Opus/Sonnet/Haiku to capture every price point

The fast tier is the biggest winner

The cheapest model available has dropped steadily over the first half of this year:

  • -Early 2026: Gemini 3 Flash at $0.50/MTok was the floor for all providers
  • -June 2026: Gemini 3.5 Flash-Lite launched at $0.25/MTok, cutting the floor in half

At $0.25/MTok input and $1.50/MTok output, Flash-Lite makes many simple tasks nearly free at scale.

What this means for your budget

Falling prices help only if your team selects the right model for each task type. Teams that default to a flagship model miss every price cut available on cheaper tier options. The savings from switching models far exceed the savings from price cuts on the same model.

  • -A 20% price cut on GPT 5.5 saves 20% on that specific model alone
  • -Switching from GPT 5.5 to GPT 5.6 Terra for summarization saves 50% on that workload
  • -Switching from Opus to Flash-Lite for classification saves 95% on input and output costs

Track which of your calls can move down a tier to capture the biggest savings available.

See your cost breakdown at /teardown or start tracking with a free dashboard.

Track your real costs

Calculators estimate. The dashboard shows what you actually spend and where you can save.