explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR: the questions people are asking
  • The prices, before and after
  • Why "20%" is conditional
  • A worked example: one long agent run
  • Does the cut help more than a model switch?
  • How to raise your cache hit rate
  • Does it change subscription value?
  • What this means for what you build
  • Caveats
  • Related reading on explainx.ai
← Back to blog

explainx / blog

Sonnet 5.5 Cache Reads Cut to $0.10: When the 20% Saving Is Real

Claude, Sonnet 5.5, Pricing, Prompt Caching, AI Costs

Claude Sonnet 5.5 cache reads fell 50% to $0.10 per million tokens. The math on when it saves 20%, when it saves 5%, and how to raise your cache hit rate.

Oct 8, 2026·9 min read·Yash Thakker
add explainx.ai
go deep
Sonnet 5.5 Cache Reads Cut to $0.10: When the 20% Saving Is Real

Anthropic halved the price of cache reads on Claude Sonnet 5.5, from $0.20 to $0.10 per million tokens, and says that makes the model "around 20% cheaper to run on most long-running work." The change was announced alongside Haiku 5.5 on October 7, 2026, and we covered the full launch in our Haiku 5.5 pricing and benchmarks post. This post is the focused version: the arithmetic behind the 20% figure, when it holds, when it does not, and what to change in your agent to capture it.

The short version: the cut is real and worth having, but "around 20%" is a best-case-ish number for cache-heavy agent loops. Your own saving depends on one thing, the share of your bill that comes from cache reads.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR: the questions people are asking

table · 2 cols
QuestionShort answer
What changed?Sonnet 5.5 cache reads: $0.20 to $0.10 per million tokens
Did input or output change?No: $2 input and $10 output per million tokens
Did cache writes change?Not that we saw: $2.50 (5 min) and $4 (1 hour)
Is it really 20% cheaper?About 20% when reads are around 94% of tokens and output about 2%; far less in chat
How big can it get?About 33% on the agentic mix SemiAnalysis uses
Do I need to do anything?No, but a higher cache hit rate means a bigger saving
Does it change my plan limits?Unknown; this is an API price change

The prices, before and after

These are Sonnet 5.5 list prices per million tokens, from our Sonnet 5.5 launch guide and the October 7 announcement.

table · 4 cols
Token typeBeforeAfterChange
Fresh input$2.00$2.00none
Cache write, 5 minutes$2.50$2.50none
Cache write, 1 hour$4.00$4.00none
Cache read$0.20$0.1050% lower
Output$10.00$10.00none

A cache read is now 5% of the fresh input price, a 95% discount. For context, caching already paid for itself quickly: with a write at $2.50 and a read at $0.10, a prompt used twice costs $2.60 cached against $4.00 uncached. Before the cut, the cached cost was $2.70.

Why "20%" is conditional

The saving on your bill depends on how much of your spend is cache reads. The cut removes $0.10 per million cache-read tokens, so:

text
saving % = 0.10 x (cache read share of tokens) / (blended price per million tokens)

The blended price is each token type's share times its price, summed. Output is the swing factor: it costs $10 per million, 100 times a cache read, so even a small share of output tokens dominates the blended price.

Here is the calculation for five token mixes, using the list prices above with the 5-minute cache write rate. These mixes are illustrative except the first two, which use the agentic and chat ratios SemiAnalysis reported in its subscription limit-testing article.

table · 4 cols
Mix (input / cache read / cache write / output)Blended beforeBlended afterSaving
SemiAnalysis agentic: 0.4% / 96.6% / 2.6% / 0.3%$0.296$0.200about 33%
Cache-heavy agent: 1% / 94% / 3% / 2%$0.483$0.389about 19.5%
Typical coding session: 2% / 90% / 4% / 4%$0.720$0.630about 12.5%
Output-heavier session: 3% / 85% / 5% / 7%$1.055$0.970about 8%
SemiAnalysis chat: 2% / 75% / 13% / 10%$1.515$1.440about 5%

Read down the table and the pattern is clear. The headline "around 20%" matches a cache-heavy agent where output is about 2% of tokens. The more your workload resembles chat, with short contexts and a lot of generation, the smaller the benefit. The more it resembles a long autonomous run that rereads a large context on every step, the bigger.

The same relationship explains a statistic from third-party cost analyses: in one maximum-effort benchmark run, cache reads were reported as about 40% of total spend. If reads are 40% of cost and the price halves, the bill drops about 20%. That is the world Anthropic's claim describes.

A worked example: one long agent run

Take an illustrative coding run of 40 steps. Each step rereads a 150,000-token cached context, adds about 2,000 fresh input tokens, and writes about 1,500 output tokens. The context was written to the cache once at the start.

table · 4 cols
Cost itemTokensBeforeAfter
Cache write, once150,000$0.375$0.375
Cache reads, 40 steps6,000,000$1.200$0.600
Fresh input80,000$0.160$0.160
Output60,000$0.600$0.600
Total$2.335$1.735

That is a saving of about 26%, because reads are half of the original bill in this run. Double the output to 120,000 tokens and the saving falls to about 20%; halve the context to 75,000 tokens and it falls to about 20% as well, since reads matter less. The run's shape, not the headline, sets your number.

Does the cut help more than a model switch?

It depends on where your tokens go. For many teams the biggest lever was never the cache read rate; it was how often the cache is hit, how long runs last, and how much output the model writes. Our analyses of Claude Code task cost on Opus 5.5 and Claude Code token efficiency and prompt cache sessions make the same point: cache behavior and effort settings move a task's bill more than a list-price tweak.

Two practical comparisons:

  • Versus Opus 5.5. Sonnet 5.5 already lists at about half Opus's input and output price. A cheaper cache read widens that gap on long runs. Whether Sonnet is the right choice still depends on task difficulty; see our Opus 5.5 versus Sonnet 5.5 comparison.
  • Versus Haiku 5.5. Haiku 5.5 lists at $0.10 input and $0.50 output, with a much cheaper cache read, so it remains the choice for subagents and high-volume simple work. See the Haiku post linked above for where each model fits.

How to raise your cache hit rate

Because the saving scales with your cache read share, the cheapest way to benefit more is to make more of your tokens cache reads. These changes cost nothing in model quality.

  1. Put stable content first. System prompt, tool definitions, project instructions and reference documents belong at the front, followed by changing content. Any change early in the prompt invalidates everything after it.
  2. Keep the prefix byte-identical. Re-sorting tools, adding a timestamp or editing a whitespace in the system prompt creates a cache miss. Generate the prefix deterministically.
  3. Reuse within the cache lifetime. A 5-minute cache expires if idle. For work with pauses, the 1-hour write at $4 may pay off; compare the extra $1.50 per million written tokens against the cost of rewriting.
  4. Do not shuffle context mid-run. Compaction, summarizing and reordering history rewrite the prefix. Do them deliberately and less often.
  5. Measure it. Log cache read, cache write and fresh input tokens per request and compute your own blended price with the formula above. If your read share is under 70%, your first win is in the prompt structure, not the price list.
  6. Control output. At $10 per million, output is the largest controllable cost in most mixes. Concise responses and an appropriate effort setting reduce it directly.

Does it change subscription value?

This is an API price change. We have no confirmation that it changes Claude plan limits or the value of a subscription. SemiAnalysis found that labs treat list prices and plan limits as separate levers: OpenAI did not change Sol token limits when it cut cache reads, which lowered API-equivalent value, while Anthropic raised Opus limits with its price cut, though not enough to offset it. If you use Sonnet 5.5 through a plan, watch your usage meter over the next week instead of assuming value moved.

What this means for what you build

If you run long-lived agents on Sonnet 5.5, you get a free discount on the part of your bill you cannot easily reduce, the cost of rereading context. Re-run your cost model with the new rate, and treat the result as the real saving, not the headline.

If you were choosing between models on cost per task, update your spreadsheet: Sonnet 5.5 improves most in cache-heavy, low-output agent work, and barely moves for chat. Pair this with prompt caching fundamentals and our guide to token pricing if you are new to the vocabulary, and the agent monthly cost breakdown for how these costs add up in a real workflow.

Caveats

  • Announcement-level facts. The $0.10 rate and the "around 20%" claim come from Anthropic's October 7 announcement as covered by our Haiku 5.5 post. We did not test billing on live accounts.
  • Our math is illustrative. The mixes are examples built from list prices. They exclude the 1.1x US-only inference premium, the 1-hour cache write rate, batch discounts and long-context surcharges. Your invoice is the truth.
  • Third-party price pages were inconsistent. Some aggregators still show $0.20 and attribute the older rate; check Anthropic's pricing page directly.

Related reading on explainx.ai

  • Claude Haiku 5.5 launch: pricing, benchmarks and the Sonnet cache cut
  • Claude Sonnet 5.5 launch and building guide
  • Claude Opus 5.5 vs Sonnet 5.5
  • What a Claude Code task costs on Opus 5.5
  • Claude Code token efficiency and prompt cache guide
  • Prompt caching and LLM cost optimization
  • SemiAnalysis: Claude vs ChatGPT subscription API value
  • AI token pricing explained

Prices are Anthropic list prices per million tokens as announced October 7, 2026, and may change. The savings table is our own illustrative arithmetic. Check Anthropic's pricing page and your invoice before making decisions.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Oct 8, 2026

Claude Max and Team API Credits: How to Claim and What They Cover

Anthropic now includes monthly Claude API credits with Max and Team plans: $100 for Max 5x, $200 for Max 20x, and a pooled balance capped at $500 for Team. Here is how to claim them, what they cover, what they do not, and the traps that will waste them.

Oct 7, 2026

Claude Haiku 5.5 Launches at $0.10 per Million Input Tokens: Pricing, Benchmarks, and When to Use It

Anthropic released Claude Haiku 5.5 on October 7, 2026: its fastest and cheapest small model, with an adjustable effort setting, computer-use scores that jump from 15.7% to 72.4%, and a 90% price cut for prompts up to 100k tokens. The same day Sonnet 5.5 cache reads got 50% cheaper and Max and Team plans gained monthly API credits.

Sep 30, 2026

GPT-6.1 Sol Launches at DevDay: Near-Astra Work at $2/$10

OpenAI launched GPT-6.1 Sol at DevDay 2026 as a capability upgrade on the same $2/$10 list price as GPT-6 Sol. The pitch is near-Astra agentic work at one-fifth of Astra's standard token prices, included on paid plans and the API the same day. This post covers the rate card, what to rerun in evals, how it sits next to Claude Sonnet 5.5 and Opus 5.5, and the one benchmark number circulating after the keynote.