July 18, 2026
Updated 2026-08-24: as published, this post priced Claude Sonnet 5 at $3.00/$15.00 and closed with an "introductory window" expiring August 31, 2026. Anthropic cancelled the increase that window pointed at, and our registry has carried the real $2.00/$10.00 rate since 2026-07-21. The Sonnet 5 figures, the routed-workload arithmetic and the closing section are corrected below; the ranking and confidence data is as published.
Twelve months ago, mid-tier models were a compromise. In July 2026, they lead the rankings in four of eight task types.
Here are the top-scoring models per task type, with their tier:
Flagships lead only on reasoning, code generation, and creative tasks. Mid-tier and fast-tier models own the rest.
Mid-tier models cost $2.00 to $3.00/MTok input versus $5.00 to $10.00 for flagships. Here is the savings breakdown:
Output savings follow the same pattern, ranging from 40% to 70% per million tokens.
Consider a team running 50,000 API calls per day across five task types. Average call uses 3,000 input and 2,000 output tokens.
All flagship (Opus 4-8):
Task-routed mid-tier (Sonnet 5 for code, Terra for QA/extraction):
Monthly savings: $49,875 by routing to mid-tier models. Quality stays within 5 confidence points.
Flagships earn their price on three workload types:
For everything else, mid-tier models deliver comparable quality at a fraction of the cost.
Claude Sonnet 5 runs at $2.00/$10.00. Of the mid-tier models compared here, that makes it the cheapest option for code-heavy workloads, and it is the top-ranked model for code review.
Anthropic announced that rate as introductory pricing through August 31, 2026, then cancelled the increase to $3.00/$15.00 that was meant to follow. There is no window to evaluate inside: the price is the price, and a team that switches keeps the saving.
Run your task-type breakdown at /teardown, or start tracking with a free dashboard.
Gemini 3.5 Flash-Lite handles classification at $0.30/MTok with 85/100 confidence, while Claude Opus 4.8 costs $5.00/MTok for lower accuracy.
We tracked every pricing change across Claude, GPT, and Gemini. The trend is clear: more cuts than increases, with new models launching at lower price points.
RAG pipelines consume 6,000+ input tokens per call. At scale, your model choice determines whether RAG costs $900/month or $18,000/month.
Calculators estimate. The dashboard shows what you actually spend and where you can save.