Anthropic halved the price of cache reads on Claude Sonnet 5.5, from $0.20 to $0.10 per million tokens, and says that makes the model "around 20% cheaper to run on most long-running work." The change was announced alongside Haiku 5.5 on October 7, 2026, and we covered the full launch in our Haiku 5.5 pricing and benchmarks post. This post is the focused version: the arithmetic behind the 20% figure, when it holds, when it does not, and what to change in your agent to capture it.
The short version: the cut is real and worth having, but "around 20%" is a best-case-ish number for cache-heavy agent loops. Your own saving depends on one thing, the share of your bill that comes from cache reads.
TL;DR: the questions people are asking
| Question | Short answer |
|---|---|
| What changed? | Sonnet 5.5 cache reads: $0.20 to $0.10 per million tokens |
| Did input or output change? | No: $2 input and $10 output per million tokens |
| Did cache writes change? | Not that we saw: $2.50 (5 min) and $4 (1 hour) |
| Is it really 20% cheaper? | About 20% when reads are around 94% of tokens and output about 2%; far less in chat |
| How big can it get? | About 33% on the agentic mix SemiAnalysis uses |
| Do I need to do anything? | No, but a higher cache hit rate means a bigger saving |
| Does it change my plan limits? | Unknown; this is an API price change |
The prices, before and after
These are Sonnet 5.5 list prices per million tokens, from our Sonnet 5.5 launch guide and the October 7 announcement.
| Token type | Before | After | Change |
|---|---|---|---|
| Fresh input | $2.00 | $2.00 | none |
| Cache write, 5 minutes | $2.50 | $2.50 | none |
| Cache write, 1 hour | $4.00 | $4.00 | none |
| Cache read | $0.20 | $0.10 | 50% lower |
| Output | $10.00 | $10.00 | none |
A cache read is now 5% of the fresh input price, a 95% discount. For context, caching already paid for itself quickly: with a write at $2.50 and a read at $0.10, a prompt used twice costs $2.60 cached against $4.00 uncached. Before the cut, the cached cost was $2.70.
Why "20%" is conditional
The saving on your bill depends on how much of your spend is cache reads. The cut removes $0.10 per million cache-read tokens, so:
saving % = 0.10 x (cache read share of tokens) / (blended price per million tokens)
The blended price is each token type's share times its price, summed. Output is the swing factor: it costs $10 per million, 100 times a cache read, so even a small share of output tokens dominates the blended price.
Here is the calculation for five token mixes, using the list prices above with the 5-minute cache write rate. These mixes are illustrative except the first two, which use the agentic and chat ratios SemiAnalysis reported in its subscription limit-testing article.
| Mix (input / cache read / cache write / output) | Blended before | Blended after | Saving |
|---|---|---|---|
| SemiAnalysis agentic: 0.4% / 96.6% / 2.6% / 0.3% | $0.296 | $0.200 | about 33% |
| Cache-heavy agent: 1% / 94% / 3% / 2% | $0.483 | $0.389 | about 19.5% |
| Typical coding session: 2% / 90% / 4% / 4% | $0.720 | $0.630 | about 12.5% |
| Output-heavier session: 3% / 85% / 5% / 7% | $1.055 | $0.970 | about 8% |
| SemiAnalysis chat: 2% / 75% / 13% / 10% | $1.515 | $1.440 | about 5% |
Read down the table and the pattern is clear. The headline "around 20%" matches a cache-heavy agent where output is about 2% of tokens. The more your workload resembles chat, with short contexts and a lot of generation, the smaller the benefit. The more it resembles a long autonomous run that rereads a large context on every step, the bigger.
The same relationship explains a statistic from third-party cost analyses: in one maximum-effort benchmark run, cache reads were reported as about 40% of total spend. If reads are 40% of cost and the price halves, the bill drops about 20%. That is the world Anthropic's claim describes.
A worked example: one long agent run
Take an illustrative coding run of 40 steps. Each step rereads a 150,000-token cached context, adds about 2,000 fresh input tokens, and writes about 1,500 output tokens. The context was written to the cache once at the start.
| Cost item | Tokens | Before | After |
|---|---|---|---|
| Cache write, once | 150,000 | $0.375 | $0.375 |
| Cache reads, 40 steps | 6,000,000 | $1.200 | $0.600 |
| Fresh input | 80,000 | $0.160 | $0.160 |
| Output | 60,000 | $0.600 | $0.600 |
| Total | $2.335 | $1.735 |
That is a saving of about 26%, because reads are half of the original bill in this run. Double the output to 120,000 tokens and the saving falls to about 20%; halve the context to 75,000 tokens and it falls to about 20% as well, since reads matter less. The run's shape, not the headline, sets your number.
Does the cut help more than a model switch?
It depends on where your tokens go. For many teams the biggest lever was never the cache read rate; it was how often the cache is hit, how long runs last, and how much output the model writes. Our analyses of Claude Code task cost on Opus 5.5 and Claude Code token efficiency and prompt cache sessions make the same point: cache behavior and effort settings move a task's bill more than a list-price tweak.
Two practical comparisons:
- Versus Opus 5.5. Sonnet 5.5 already lists at about half Opus's input and output price. A cheaper cache read widens that gap on long runs. Whether Sonnet is the right choice still depends on task difficulty; see our Opus 5.5 versus Sonnet 5.5 comparison.
- Versus Haiku 5.5. Haiku 5.5 lists at $0.10 input and $0.50 output, with a much cheaper cache read, so it remains the choice for subagents and high-volume simple work. See the Haiku post linked above for where each model fits.
How to raise your cache hit rate
Because the saving scales with your cache read share, the cheapest way to benefit more is to make more of your tokens cache reads. These changes cost nothing in model quality.
- Put stable content first. System prompt, tool definitions, project instructions and reference documents belong at the front, followed by changing content. Any change early in the prompt invalidates everything after it.
- Keep the prefix byte-identical. Re-sorting tools, adding a timestamp or editing a whitespace in the system prompt creates a cache miss. Generate the prefix deterministically.
- Reuse within the cache lifetime. A 5-minute cache expires if idle. For work with pauses, the 1-hour write at $4 may pay off; compare the extra $1.50 per million written tokens against the cost of rewriting.
- Do not shuffle context mid-run. Compaction, summarizing and reordering history rewrite the prefix. Do them deliberately and less often.
- Measure it. Log cache read, cache write and fresh input tokens per request and compute your own blended price with the formula above. If your read share is under 70%, your first win is in the prompt structure, not the price list.
- Control output. At $10 per million, output is the largest controllable cost in most mixes. Concise responses and an appropriate effort setting reduce it directly.
Does it change subscription value?
This is an API price change. We have no confirmation that it changes Claude plan limits or the value of a subscription. SemiAnalysis found that labs treat list prices and plan limits as separate levers: OpenAI did not change Sol token limits when it cut cache reads, which lowered API-equivalent value, while Anthropic raised Opus limits with its price cut, though not enough to offset it. If you use Sonnet 5.5 through a plan, watch your usage meter over the next week instead of assuming value moved.
What this means for what you build
If you run long-lived agents on Sonnet 5.5, you get a free discount on the part of your bill you cannot easily reduce, the cost of rereading context. Re-run your cost model with the new rate, and treat the result as the real saving, not the headline.
If you were choosing between models on cost per task, update your spreadsheet: Sonnet 5.5 improves most in cache-heavy, low-output agent work, and barely moves for chat. Pair this with prompt caching fundamentals and our guide to token pricing if you are new to the vocabulary, and the agent monthly cost breakdown for how these costs add up in a real workflow.
Caveats
- Announcement-level facts. The $0.10 rate and the "around 20%" claim come from Anthropic's October 7 announcement as covered by our Haiku 5.5 post. We did not test billing on live accounts.
- Our math is illustrative. The mixes are examples built from list prices. They exclude the 1.1x US-only inference premium, the 1-hour cache write rate, batch discounts and long-context surcharges. Your invoice is the truth.
- Third-party price pages were inconsistent. Some aggregators still show $0.20 and attribute the older rate; check Anthropic's pricing page directly.
Related reading on explainx.ai
- Claude Haiku 5.5 launch: pricing, benchmarks and the Sonnet cache cut
- Claude Sonnet 5.5 launch and building guide
- Claude Opus 5.5 vs Sonnet 5.5
- What a Claude Code task costs on Opus 5.5
- Claude Code token efficiency and prompt cache guide
- Prompt caching and LLM cost optimization
- SemiAnalysis: Claude vs ChatGPT subscription API value
- AI token pricing explained
Prices are Anthropic list prices per million tokens as announced October 7, 2026, and may change. The savings table is our own illustrative arithmetic. Check Anthropic's pricing page and your invoice before making decisions.
