Editorial4 min read

Cache Discounts: Which Provider Saves You the Most?

July 14, 2026

Cache pricing varies more than you think

Every provider offers discounts on cached input tokens. The size of that discount changes which provider is cheapest for your workload.

Here are the cache read discount multipliers from the model registry:

  • -Anthropic (all models): 0.10 (90% off list price)
  • -Google (all models): 0.25 (75% off list price)
  • -OpenAI (all models): 0.10 (90% off list price)

Anthropic and OpenAI tie at 90% off. Google trails at 75% off on cached reads.

Cache write surcharges

Writing to cache costs extra. Anthropic and OpenAI charge a write premium:

  • -Anthropic: 1.25x (25% surcharge on input)
  • -Google: 1.0x (no surcharge)
  • -OpenAI: 1.25x (25% surcharge on input)

Google charges zero extra for cache writes. For workloads with low cache reuse, Google saves money on the write side.

How caching changes the leaderboard

Consider Gemini 3.1 Pro versus Claude Sonnet 5 for a summarization workload. Without caching, at 10,000 input tokens per call:

  • -Gemini 3.1 Pro: 10,000/1M x $2.00 = $0.020/call input
  • -Claude Sonnet 5: 10,000/1M x $3.00 = $0.030/call input

Gemini is 33% cheaper. Now apply 80% cache hit rate:

  • -Gemini cached: (2,000/1M x $2.00) + (8,000/1M x $2.00 x 0.25) = $0.008/call
  • -Sonnet 5 cached: (2,000/1M x $3.00) + (8,000/1M x $3.00 x 0.10) = $0.0084/call

With 80% cache hits, the gap narrows from 33% to just 5%. At 90% cache hits, Sonnet 5 pulls ahead.

The breakeven cache rate

For Sonnet 5 to beat Gemini 3.1 Pro on input cost, you need roughly 85% cache hit rates. Below that threshold, Gemini wins on price. Above it, Anthropic's deeper discount tips the balance.

Flagship tier caching comparison

At $5.00/MTok input for both Opus and GPT 5.5, the 90% discount brings cached input to $0.50/MTok. The tie holds even after caching.

Output pricing breaks that tie. Opus at $25.00/MTok output is 17% cheaper than GPT 5.5 at $30.00/MTok.

Three rules for cache optimization

1. Measure your cache hit rate first. The discount only matters if your hits are above 50%. 2. System prompts are free wins. Identical system prompts across calls cache at near 100% rates. 3. Provider choice depends on hit rate. Below 85% hits, Google's lower list prices often win. Above 85%, Anthropic's deeper discount dominates.

Model your cache savings at /teardown, or start tracking cache hit rates with a free dashboard.

Track your real costs

Calculators estimate. The dashboard shows what you actually spend and where you can save.