July 1, 2026
We track every pricing change across all three major providers in a single registry. The data from the first half of 2026 reveals a clear and consistent downward trend.
Six downward moves versus one upward move tell a clear story about market direction. New models launch at price points that undercut or match their predecessors across all three providers.
The one increase, Opus moving from $4.50 to $5.00, likely reflects Anthropic repositioning it as premium. That repositioning happened just ahead of the Sonnet 5 launch at the mid-tier price point.
The cheapest model available has dropped steadily over the first half of this year:
At $0.25/MTok input and $1.50/MTok output, Flash-Lite makes many simple tasks nearly free at scale.
Falling prices help only if your team selects the right model for each task type. Teams that default to a flagship model miss every price cut available on cheaper tier options. The savings from switching models far exceed the savings from price cuts on the same model.
Track which of your calls can move down a tier to capture the biggest savings available.
See your cost breakdown at /teardown or start tracking with a free dashboard.
Gemini 3.5 Flash-Lite handles classification at $0.25/MTok with 85/100 confidence, while Claude Opus costs $5.00/MTok for lower accuracy.
RAG pipelines consume 6,000+ input tokens per call. At scale, your model choice determines whether RAG costs $900/month or $18,000/month.
Anthropic offers 90% off cached reads. Google offers 75%. OpenAI matches at 90%. Here is how cache discounts shift the pricing landscape.
Calculators estimate. The dashboard shows what you actually spend and where you can save.