Intelligence
Mid-tier models now score 90+ confidence on most tasks at 40-60% of flagship pricing. The data shows flagships are overkill for common workloads.
A snapshot of current pricing across Claude, GPT, and Gemini, plus which models deliver the best value.
Anthropic offers 90% off cached reads. Google offers 75%. OpenAI matches at 90%. Here is how cache discounts shift the pricing landscape.
Gemini 3.5 Flash-Lite handles classification at $0.25/MTok with 85/100 confidence, while Claude Opus costs $5.00/MTok for lower accuracy.
OpenAI released the GPT 5.6 family on July 9, 2026. Three tiers cover flagship, mid, and fast, all with 1M context windows.
RAG pipelines consume 6,000+ input tokens per call. At scale, your model choice determines whether RAG costs $900/month or $18,000/month.
We tracked every pricing change across Claude, GPT, and Gemini. The trend is clear: more cuts than increases, with new models launching at lower price points.
Anthropic released Claude Sonnet 5 on June 30 at $3/$15 per MTok, with introductory pricing of $2/$10 through August 31, 2026.
June 2026 saw two major launches: Claude Fable 5 and Claude Sonnet 5. Google cut Gemini Flash output pricing. The mid tier became the new battleground.
Anthropic launched Claude Fable 5 on June 9, 2026 at $10.00 input and $50.00 output per million tokens. It is the most expensive model in the registry.
Gemini 3.5 Flash output pricing fell from $0.75 to $0.60 per million tokens in June 2026, widening Google's cost advantage in the fast tier.
Anthropic raised Claude Opus 4.8 input pricing from $4.50 to $5.00 per million tokens on May 28, 2026. Here is what it means for flagship users.
GPT-5.5 input cost dropped from $7.50 to $6.00 per million tokens in May 2026. What it means for teams running flagship workloads.
Get per-call cost visibility across Claude, GPT, and Gemini. Free tier tracks up to 1,000 calls/month.