explainx.ai0k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • What should you check before you trust a percent-off headline?
  • What does the same token mix cost on Opus 5 and Opus 5.5?
  • How much does the cache hit rate change one task?
  • Why does one output token match one hundred cache reads?
  • Does a higher effort level pay for itself?
  • Which model should do which part of the task?
  • What does a cache write cost, and what misses the cache?
  • When does /compact pay for itself?
  • Does fast mode or the Batch API change an interactive task?
  • How do you check a real task instead of an illustration?
  • What did one prompt audit change, and what should you ignore?
  • What this guide will not decide for you
  • Related on explainx.ai
← Back to blog

explainx / blog

What a Claude Code Task Costs on Opus 5.5

Claude Code, Opus 5.5, Token Economics, Prompt Caching, Guides

A same-token Opus 5.5 task costs about $2.40 versus $3.50 on Opus 5. Cache, turns, and effort move the Claude Code bill more than the sticker.

Sep 27, 2026·19 min read·Yash Thakker
add explainx.ai
go deep
What a Claude Code Task Costs on Opus 5.5

Anthropic's published estimate is that typical token-billed workloads cost "40% less to run" on Opus 5.5 at default settings. That sentence, from Addy Osmani's September 25, 2026 post What a task costs on Opus 5.5 (also on claude.dev), stacks two changes. List prices fell. On top of that, Anthropic assumes Opus 5.5 spends fewer tokens per task because its default effort is medium, where Opus 5's default was high. The 40% line is an estimate of a finished task under those defaults. It is a poor description of the price of one token.

Hold the token mix fixed and the price cut is smaller. An illustration with 2.0 million cache reads, 200,000 fresh input tokens, and 60,000 output tokens costs $3.50 on Opus 5 and $2.40 on Opus 5.5, about 31%. The list prices themselves live on Anthropic's Opus 5.5 model overview: input $4 per million tokens, output $20, a 5-minute cache write $5, a 1-hour cache write $8, and a cache read $0.20. explainx.ai's launch pricing post already records those stickers. This guide answers the next question: given those rates, what does one Claude Code task cost, and which session habits move the receipt more than the sticker.

What should you check before you trust a percent-off headline?

table · 2 cols
QuestionDirect answer
Is a token 40% cheaper?Input and output fell 20% ($5 to $4, $25 to $20). Cache reads fell 60% ($0.50 to $0.20). The 40% line adds an assumption of fewer tokens at medium effort.
What does one same-token task cost?Illustration: $2.40 on Opus 5.5 versus $3.50 on Opus 5.
What does that look like for a month?Ten such tasks a working day, 22 days: about $528 versus $770. On Pro, Max, and Team this is a limit, not an invoice.
Which line dominates a warm session?Cache reads, at $0.20 per million — 5% of a fresh input token, down from 10% on Opus 5.
What costs as much as 60,000 output tokens?$1.20, the same as 6 million cache reads. Thinking bills as output.
What is the default effort?Medium. Opus 5's default was high. Levels are not the same amount of thinking across the two models.
How do I read the receipt?claude update, then /usage. The Session block is the task. /cost is the same command.
What misses the cache?A pause longer than the lifetime, a model switch, compaction, connecting or disconnecting an MCP server, the first turn of fast mode, and an effort change on Bedrock, Agent Platform, or a gateway.
text
/usage
/effort medium
/effort status
/model
/compact
/clear

Run /usage on the task you just finished. The three ratios worth writing down are cache reads divided by total input, output divided by input, and total input compared with the size of the conversation. A huge input total on a modest window means the prefix was re-sent on every turn. That is the whole economic story of a coding agent, which is also why agent loops cost more than chat even when the model sticker looks modest.

What does the same token mix cost on Opus 5 and Opus 5.5?

This table is an illustration using published list rates. It is not a log from an explainx.ai session. The mix is the one Osmani's post uses to separate the price cut from the "fewer tokens" assumption: 2.0 million cache reads, 200,000 fresh input tokens, 60,000 output tokens.

table · 6 cols
LineTokensOpus 5 rateOpus 5Opus 5.5 rateOpus 5.5
Cache reads2.0M$0.50 / MTok$1.00$0.20 / MTok$0.40
Fresh input200K$5 / MTok$1.00$4 / MTok$0.80
Output60K$25 / MTok$1.50$20 / MTok$1.20
Total$3.50$2.40

The $1.10 gap is 31% of $3.50. Cache reads contribute $0.60 of that gap, fresh input $0.20, and output $0.30. Most of the same-token savings sits on the read line, because the read rate went from one tenth of input to one twentieth. A cached token on Opus 5.5 costs 5% of a fresh input token.

Ten illustrations a day, 22 working days:

table · 3 cols
Per taskPer month (220 tasks)
Opus 5$3.50$770
Opus 5.5$2.40$528
Difference$1.10$242

That $242 is the price cut alone, on this mix. Anyone quoting 40% on the same tokens is mixing in Anthropic's second assumption: Opus 5.5 finishes the work with fewer tokens at medium effort. Measure that on your repo. Do not import it from the headline.

On Claude Pro, Max, and Team, this table is not a bill. Anthropic passes the lower price through to usage limits, about 25% further once cached context is included, and five-hour limits also went up. A limit reset lives under Settings > Usage on the web app or in Claude Desktop. It does not live in the Claude Code terminal. The reset applies to the account, Claude Code included. If you are on a plan, /usage still tells you how hard you are pushing the window. It does not print an invoice.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

How much does the cache hit rate change one task?

Agent loop beside a short chat, with the meter rising on the loop: the shape of a Claude Code task once cache reads, not the sticker, dominate the bill

A second illustration, still at Opus 5.5 input rates only. Context grows from 20,000 to 120,000 tokens across 40 turns. The average turn is about 70,000 tokens, so input sums to about 2.8 million tokens (40 × 70,000). Output is left out of this table on purpose, so the cache effect is visible by itself.

table · 4 cols
Cache share of those 2.8M input tokensCache-read tokensFresh inputInput cost
0% (every token billed as fresh input)02.8M$11.20
90%2.52M280K$1.62
96%2.688M112K$0.99

The arithmetic at 90%: 2.52 million reads × $0.20 is $0.504, and 0.28 million fresh tokens × $4 is $1.12, which sums to $1.62. At 96% the same sum is about $0.99. The uncached case is 2.8 × $4 = $11.20. Moving from a cold transcript to a 90% hit rate cuts input cost by roughly seven times. Moving from 90% to 96% saves about $0.63 on this illustration. The first jump is the one that changes whether Opus is affordable. The second jump is real money, and it is smaller than people expect once the cache is already warm.

The same task in 25 turns, at the same average of about 70,000 tokens, is about 1.75 million input tokens. At a 90% hit rate that input costs about $1.02 (1.575 million reads × $0.20 = $0.315, plus 0.175 million fresh × $4 = $0.70). Fifteen fewer round trips save about $0.60 of input on this sketch, before you count the output those extra turns would have produced.

A warm turn still scales with context size, because each turn re-reads the prefix:

table · 3 cols
Cached context on one turnOne cache read30 such turns
20,000 tokens$0.004$0.12
150,000 tokens$0.030$0.90

Thirty turns at 150,000 cached tokens are about $0.90 of reads. Thirty turns at 20,000 are about $0.12. The model sticker did not change between those rows. The window did. That is why a habit that keeps the prefix short often saves more than waiting for the next price cut.

Why does one output token match one hundred cache reads?

Output is $20 per million. A cache read is $0.20 per million. The ratio is 100×. Sixty thousand output tokens cost $1.20, and so do six million cache reads. In the same-token receipt above, output is already half of the $2.40 ($1.20 of $2.40). A session that looks "cheap on input" because the cache is hot can still be an output bill.

Thinking is billed as output even when Claude Code shows you a summary of the reasoning. You pay for tokens you do not see in full. That is the mechanical reason effort levels show up in /usage as output, and it is why a retry is expensive twice: you pay the failed attempt's output, then you pay the context of the retry, which now includes the failed attempt unless you rewind.

Does a higher effort level pay for itself?

Opus 5.5 effort levels are low, medium, high, xhigh, plus max for a single session. The default is medium. Opus 5's default was high. A named level is not a fixed thinking budget shared across models. At a given level, Opus 5.5 thinks more per turn than Opus 5, and the gap is widest at xhigh and max. Treat the ladder as a control on this model, then re-check /usage if you switch models. The prompting guide is the companion on what to put in the prompt. This section is only the price of the thinking the prompt triggers. The Shift+Tab effort proposal is a separate product question about whether plan mode survives. The dollars below apply either way.

An illustration: 20,000 extra thinking tokens across a task, billed as output, is 0.02 × $20 = $0.40. A 10-turn retry sitting on 100,000 cached tokens of context, plus 10,000 output tokens, lands in the same neighborhood: 1 million cache reads × $0.20 = $0.20, and 10,000 output × $20 per million = $0.20, about $0.40 together. High effort pays for itself when it removes one retry. It is wasted when medium would have finished.

text
/effort medium
/effort status
/model

/effort status tells you the level on the current session. /model is where the default is saved, so a one-off switch back to Opus 5 or over to Sonnet will still be there tomorrow unless you switch it back.

Changing effort keeps the cache on an API key or a Claude subscription. On Amazon Bedrock, Google Cloud's Agent Platform, or a Claude apps gateway, an effort change clears the cache, and the next turn pays a write on the whole conversation. Thinking cannot be turned off on Opus 5.5. Low is the floor, not a return to a model that answers with no reasoning tokens.

Which model should do which part of the task?

Price the job, then pick the model. The ladder explainx.ai uses when reading Osmani's routing notes:

table · 3 cols
JobModelWhy the rate matters
Search and log subagentsSonnet or HaikuThe parent session should not hold the file dump.
Supervised feature work, debugging, review that editsOpus 5.5You are in the loop. Cache reads at $0.20 dominate once the prefix is warm.
Unsupervised long runs, no existing pattern, many subagentsFable 5.1The result matters more than the token price.

Fable 5.1 lists at $10 / $50 per million input and output tokens. Its cache reads are $0.25 per million: 0.025× its own input price, and only 1.25× Opus 5.5's cache-read rate ($0.25 / $0.20). The headline gap is the fresh input and the output, not the warm read. Our Fable 5.1 versus Opus 5.5 comparison is the quality side of that choice. If the alternative you are weighing is GPT-6 Sol, the Sol versus Opus 5.5 comparison is a different sticker, and it still will not tell you your cache share. /usage will.

Switch models at a break in the work. The first turn on the new model pays the cache write on the whole conversation. Run /compact first, or start a fresh session and paste a short plan, so that write lands on a small prefix. /model saves the default. Switch back when the hard stretch is over, or tomorrow's search subagents will still be on Fable.

The opusplan alias plans with Opus and edits with Sonnet. That puts the code edits on Sonnet. It is the opposite of "keep edits on Opus 5.5 and send lookups to a smaller model." Measure a real task with /usage before you make opusplan the default. If your edits are the part you wanted Opus for, the alias is solving a different problem.

Subagent model, in the agent definition:

yaml
model: haiku

model: sonnet is the same knob. CLAUDE_CODE_SUBAGENT_MODEL sets a default for subagents. A model named in the definition overrides that environment variable. Agent teams are experimental. In plan mode they run at about seven times the tokens of a standard session. Turn that on when you want the parallel exploration, and read /usage after the first one so the multiplier is yours, not a blog's.

What does a cache write cost, and what misses the cache?

Claude Code's cache lifetime is one hour on a subscription and five minutes on an API key or a cloud provider. A subscription that has started drawing usage credits also falls back to five minutes. Each hit resets the lifetime. Writes bill at 1.25× input for the five-minute cache ($5 per million on Opus 5.5) and 2× input for the one-hour cache ($8 per million).

At 120,000 tokens, an illustration:

table · 3 cols
EventRateCost
5-minute cache write$5 / MTok$0.60
1-hour cache write$8 / MTok$0.96
Cache read$0.20 / MTok$0.024

One five-minute write at this size costs about as much as 25 reads ($0.60 / $0.024). A six-minute gap on an API key turns the next turn's read into a write. The session did not get smarter in that minute. The bill did. A one-hour write at the same 120,000 tokens is about $0.96, so a subscription user who steps away for lunch and comes back inside the hour is in a different regime from an API-key user who steps away for a standup.

These actions miss the cache:

  • A pause longer than the lifetime.
  • An effort change on Bedrock, Agent Platform, or a Claude apps gateway.
  • The first request after you turn fast mode on.
  • Connecting or disconnecting an MCP server.
  • Switching models.
  • Compaction.

On an API key or a Claude subscription, an effort change by itself does not miss. People who learned the older "any effort change busts the cache" rule should re-check which product they are billed through. The miss list above is the one that matches the September 25, 2026 task-cost notes.

When does /compact pay for itself?

/compact at 150,000 tokens with a warm cache is about $0.25 in the published illustration, and it pays back in about ten later turns if each later turn saves about $0.025 of cache reads. After a cold cache, that same compact's input alone is about $0.75 on a five-minute cache, because 150,000 tokens at the $5 write rate is $0.75 before the summary is even generated. Compact before a break, while the prefix is still warm. Compacting after the lifetime expired means you pay a write to discover you should have written a shorter prompt an hour ago.

/clear between unrelated tasks. A new task that inherits a 150,000-token prefix pays the read on every turn for context that cannot help it. /rewind, while the cache is still warm, returns you to a cached prefix, so you drop the bad attempt without a full rewrite of the conversation. Costs-oriented docs suggest keeping CLAUDE.md under 200 lines. MCP tool definitions are deferred: Claude Code starts with names and server instructions, and loads the full schema when a tool is actually needed. A long CLAUDE.md plus a dozen always-on tool schemas is a tax on every turn, cached or not, because it raises the prefix those $0.004-versus-$0.030 rows are measuring.

Does fast mode or the Batch API change an interactive task?

Fast mode is up to 2.5× faster at 2× price: $8 input and $40 output per million tokens on Opus 5.5. On a subscription it bills usage credits, not plan limits. The first request after you turn it on pays fast-mode input on the whole conversation, and that prefix is uncached. Turn it on at the start of the session you want to be fast. Turning it on at turn 30 of a 150,000-token thread spends the expensive rate on history you have already paid to cache at the normal rate.

An illustration of that first fast-mode turn at 120,000 tokens of input is 0.12 × $8 = $0.96 of input, before output. At the normal cache-read rate the same prefix is $0.024. The speed is real. The moment you flip the switch is the moment the receipt changes shape.

The Batch API is half price on input and output ($2 and $10 per million at these Opus 5.5 list rates). Interactive Claude Code is not a batch job. Do not expect /usage in a live session to show the batch discount. Batch matters when you are running offline evals or a queue of similar tasks through the API, not when you are pair-programming in the terminal.

Anthropic's public enterprise costs page puts average spend at about $13 per developer per active day, with 90% of users under $30 per active day, on current models. Those are fleet averages, not a promise about your repo. A single Opus 5.5 task at the $2.40 illustration is well under that daily band. A cold 2.8 million-token input at $11.20, plus output, plus a retry, can clear a large fraction of the $30 band in one afternoon. The average is a backdrop. /usage is the measurement.

How do you check a real task instead of an illustration?

Install or refresh Claude Code so you are on v2.1.280 or later:

text
claude update

Then, inside the session:

text
/usage

Read the Session block. Run the same real task once on Opus 5 and once on Opus 5.5 if you still have both available, and compare three things:

  1. Cache share. Cache-read tokens divided by total input tokens. Under 50% on a long task usually means a miss: a pause, a model switch, a compact at the wrong time, or a provider that cleared the cache.
  2. Output versus input. If output dollars exceed input dollars, effort, thinking, and retries are the lever. Another 20% off the input sticker will not fix that.
  3. Total input versus conversation size. Total input should be much larger than the final window, because every turn re-sends the prefix. If total input is only a little larger than the window, you had very few turns, or the cache accounting is not what you think it is. If total input is enormous relative to a short final window, you took many turns while the window was already large.

The illustrations in this post exist so those three numbers have a scale. They are not targets. A 40-turn task in a strange repo will not match 2.8 million input tokens. Your Session block will.

What did one prompt audit change, and what should you ignore?

Osmani's post describes /claude-api prompt-audit and one internal support benchmark: 44 tickets, moving from Opus 4.8 to Opus 5.5 at low effort, which cut cost by about 18%. The audit then cut a further 9%, landing about 25% below the Opus 4.8 starting point. The cuts came from removing a mandatory six-step procedure, a scratchpad rule, a verify-twice rule, and contradictory instructions.

That is one benchmark. It is not a number to budget. The useful part is the shape of the waste: instructions that force extra steps, extra notes, and a second verification pass all show up as output tokens, and output is the 100× line. If your CLAUDE.md tells the model to restate the plan, keep a scratchpad in the reply, and then verify the scratchpad, you are buying output on purpose. Audit the instructions against the Session block. If low effort plus a shorter prompt finishes the ticket, the six-step ritual was a cost center.

text
/claude-api prompt-audit

What this guide will not decide for you

The $2.40 and $1.62 figures are illustrations at published rates. Your cache share, your output length, and your retry rate dominate them. A plan subscriber never sees these dollars as a charge. The limit moves, and the place to see a reset is Settings, not the terminal.

Effort names do not transfer cleanly from Opus 5. High on Opus 5.5 thinks more than high on Opus 5. Copying a teammate's /effort high habit without reading /usage imports their quality preference and their output bill.

opusplan and agent teams can raise spend while feeling like optimizations. Edits on Sonnet are cheaper per token and are the wrong default if you chose Opus so the edits would be Opus. A seven-times token multiplier in an experimental team mode needs one measured session before it becomes a standing workflow.

Fast mode's first turn is the expensive one. Batch pricing does not apply to the session you are typing in. And the 44-ticket audit is a single internal result. Quote it as a case, then measure your own prompts.

The sticker is settled: 20% off fresh input and output, 60% off cache reads, and a same-token illustration about 31% cheaper. The session is not settled until you read the Session block. Start with /usage on the next real task, keep medium unless a retry is already costing you the $0.40, and compact while the cache is warm.

Related on explainx.ai

  • Claude Opus 5.5 launch: benchmarks and list prices
  • Opus 5.5 prompting guide
  • Fable 5.1 versus Opus 5.5
  • GPT-6 Sol versus Opus 5.5
  • Plan mode, Shift+Tab, and effort levels
  • Why agent loops cost more than chat
  • Addy Osmani, What a task costs on Opus 5.5 (September 25, 2026)
  • Same post on claude.dev
  • Opus 5.5 model overview (list prices)

List prices, cache lifetimes, and Claude Code commands in this guide match published figures as of September 27, 2026. Dollar rows marked as illustrations use those rates on assumed token mixes. They are not invoices from a live explainx.ai session. Re-check the model overview and /usage before you budget a quarter on them.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 23, 2026

How to Actually Use Claude Opus 5.5: Anthropic's Own Prompting Playbook

Anthropic published a developer playbook the same day Opus 5.5 launched, and it contains some genuinely counter-intuitive advice — stop telling the model to "think carefully" (it always does now), hand over entire tasks instead of micromanaging steps, and when a design comes out generic, list the specific patterns you don't want rather than asking for something vaguely "not generic." Here's the full guide, condensed.

Jul 13, 2026

Claude Code vs OpenCode Token Overhead — What Systima Measured at the API Boundary

HN hit 456 points on harness overhead, not model IQ. Systima spliced a proxy between Claude Code 2.1.207 and OpenCode 1.17.18 — same model, same machine. explainx.ai maps the floor, multipliers, cache economics, and what to do about it.

Sep 27, 2026

How to Make an Opus 5.5 Video in Claude Code

The viral Opus 5.5 films are not a video model. This is the practical path: install John Heibel's starter kit, ask Claude Code for a 15-second cartoon, review the storyboard and contact sheets, then render the MP4.