September 29, 2026 — A week after Opus 5.5 became the Claude Code default, Sonnet 5.5 arrived at half the list price and, on Anthropic’s own Terminal-Bench 4.0 row, ahead of the flagship. Slack then did the predictable thing: “swap the default.”
That is the wrong read of the table. The models converged on knowledge-work Elo. They did not become interchangeable. Artificial Analysis still has Opus 5.5 (max) at 58 and Sonnet 5.5 (max, default fallback) at 56. At that max setting, Sonnet cost more per index task than Opus. Third-party code review and Terminal-Bench reruns do not all agree with Anthropic’s 70.6% headline.
This comparison is the routing doc: when the $2/$10 sticker is real savings, and when “just turn Max on” is how you overpay.
TL;DR — the questions teams ask after the screenshot
| Question | Direct answer |
|---|---|
| Who is the flagship? | Opus 5.5. Claude Code default. AA Intelligence Index 58. |
| Who is the everyday 5.5? | Sonnet 5.5 at $2/$10, native 1M, thinking on. Alias sonnet = medium in Claude Code v2.1.284+. |
| Who wins Anthropic Terminal-Bench 4.0? | Sonnet 70.6% vs Opus 66.4% (Xhigh). |
| Who wins AA Intelligence Index? | Opus 58 vs Sonnet 56. |
| Cheaper per million tokens? | Sonnet, always, on the rate card. |
| Cheaper per AA max task? | Opus (~$5.98) vs Sonnet (~$7.60). |
| Cheaper at matched score 56? | Opus Xhigh (~$3.46) vs Sonnet max (~$7.60). |
| Default swap? | No. Route by scope, not by the Terminal-Bench cell. |
For OpenAI’s flagship against Sonnet, use Sonnet 5.5 vs GPT-6 Astra. For OpenAI’s mid-tier against Opus, use GPT-6 Sol vs Opus 5.5.
Anthropic’s launch table (read the footnotes)

Source: Anthropic Sonnet 5.5 launch evals (System Card methodology), September 28, 2026. Same table as the building guide. GPT-6 Sol is a mid-tier OpenAI column, not Astra.
| Eval | Sonnet 5.5 | Opus 5.5 | Gap |
|---|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 66.4% (Xhigh) | Sonnet +4.2 |
| FrontierCode 1.1 Main | 52.1% Xhigh / 46.2% Max | 54.4% | Opus, and Sonnet loses at Max |
| CursorBench 4.0 | 55.5% | 57.8% | Opus +2.3 |
| GDPval-AA v2.1 (Elo) | 1844 | 1846 | Tie for practical purposes |
| AA-Briefcase v1.1 (Elo) | 1811 | 1822 | Opus +11 Elo |
| HLE with tools | 64.5% | 67.7% | Opus +3.2 |
| OSWorld 2.1 (partial) | 80.1% | 81.8% | Opus +1.7 |
| Chartography (no tools) | 61.6% | 64.4% | Opus +2.8 |
What the table is allowed to mean: for well-specified terminal / coding agents, Sonnet 5.5 is in the same band as Opus and wins the vendor Terminal-Bench row at half the sticker. For merge-clean agentic coding (FrontierCode), IDE agent grading (CursorBench), hard exams, computer use, and charts, Opus still leads. Knowledge-work Elo is not a reason to fire Opus; it is a reason to stop treating Sonnet as a toy.
Four footnotes that rewrite Slack threads:
- Opus Terminal-Bench is Xhigh, not “default Opus.”
- Sonnet FrontierCode Max (46.2%) is worse than Xhigh (52.1%). FrontierCode penalizes diffs that would not merge without human edits. At Max, Claude Code’s code-review skill fans work across subagents; Cognition saw timeouts and extra out-of-scope edits. Max is not a free quality slider.
- Artificial Analysis ran some knowledge-work evals on a pre-release Claude Platform build with a structured-output bug. Anthropic says it is fixed and any leftover effect understates Sonnet.
- GPT-6 Sol Chartography / some AA rows may predate an OpenAI image-understanding fix. That column is not the Astra comparison.
Effort-cost plots for Briefcase, CursorBench, FrontierCode, and Terminal-Bench live in the Sonnet 5.5 building guide — they are the visual version of “Low/Med Sonnet often matches old Sonnet 5 at a fraction of the dollars.”
Artificial Analysis: two points, and the invoice flips
AA’s Sonnet 5.5 article (Intelligence Index v4.3.2, ten evals, adaptive reasoning, default fallback):
- Opus 5.5 max: 58 at about $5.98 per index task; ~119k / 84k input/output tokens per task in AA’s Opus max row; 260M output tokens across the suite.
- Sonnet 5.5 max: 56 at about $7.60 per index task; ~193k output tokens per task, highest AA has measured; about 60% more output than Opus max; 410M output tokens across the suite.
- Terminal-Bench 4.0 (AA): Sonnet about 64% vs Opus (and Astra xhigh) about 60%. Same direction as Anthropic, different absolute scores.
- AA-Briefcase / GDPval-AA / AutomationBench-AA: near parity, “albeit with significantly higher token usage.”
- AA-Omniscience: Sonnet 54% factual accuracy vs Opus 66%, with a lower hallucination rate (47% vs 59%). Smaller model, fewer confident wrong facts, also fewer facts.
- HLE and SciCode: Sonnet about 6 points behind Opus on AA’s write-up.
- Terminal-Bench-Science (not in the index): Sonnet 53%, behind Opus and GPT-6 Astra.
AA’s published effort ladders (same fallback setting):
| Effort | Sonnet 5.5 index | Sonnet ~$/task | Opus 5.5 index | Opus ~$/task |
|---|---|---|---|---|
| Low | 36 | $0.41 | 42 | $0.55 |
| Medium | 41 | $0.59 | 51 | $1.34 |
| High | 47 | $1.08 | 54 | $1.82 |
| Xhigh | 52 | $2.74 | 56 | $3.46 |
| Max | 56 | $7.60 | 58 | $5.98 |
The matched-score trap: Sonnet max and Opus Xhigh both sit on 56. Sonnet’s AA bill at that point is more than 2× Opus Xhigh ($7.60 vs $3.46). Half-price tokens do not survive max verbosity.
Where Sonnet actually saves on AA: low and high (and typically medium) — cheaper dollars for a lower score. That is the intended product: volume at medium/high, not max as a fake Opus.
AA also says Sonnet 5.5 sits off the Intelligence vs cost-per-task Pareto frontier at max, and that high is the most competitive Sonnet effort on that plot (narrowly behind some GPT-6 Sol configs at similar $/task). Treat that as their mix, not your mix.
Fallback rate on the index: about 0.1% of Sonnet tasks, mostly Terminal-Bench, falling back to Sonnet 5.
Other independent scoreboards (do not average them)
Vals AI Index (Max effort on most tests, server-side fallbacks allowed): Sonnet 69.22% ±0.96 vs Opus 69.69% ±0.94 — gap inside the error bars. Cost per Vals test: about $20.80 Sonnet vs $32.77 Opus, so Vals’ mix still has Sonnet cheaper, the opposite of AA’s max-index mix. Vals Terminal-Bench 4.0: Sonnet 53.03% vs Opus 61.62%. Counting fallback-assisted tasks as failures: Sonnet 50.51%, Opus 53.54%. That is Opus ahead on terminal, Anthropic and AA have Sonnet ahead. Three labs, three stories. Your harness is a fourth.
Code review (CodeRabbit, 13 hardest known-bug cases): Sonnet 5.5 caught 6/13 with actionable comments vs 4/13 for Sonnet 5, similar precision, about half the wall-clock, and list-price Claude calls about 60% cheaper per review than Sonnet 5. Opus 5.5 caught 8 (standard) to 10 (max) of the same 13 at higher precision. Hard review is open-ended judgment — exactly where Anthropic still tells you to keep Opus. Sonnet 5.5 is the fast PR bot; Opus is the high-risk diff.
AWS Bedrock pairing note (September 28, 2026): Opus for release debugging, multi-PR stacks, security review of large PRs, migrations, long analyses, contract redlining. Sonnet where the approach is already clear and you need speed and lower cost per task for most work — alerts, SQL, UI tests, capped IDE agents, short spreadsheet edits. That is the same split as Addy Osmani’s claude.dev building guide.
List prices, cache, and Claude Code reality
| Opus 5.5 | Sonnet 5.5 | |
|---|---|---|
| Model ID | claude-opus-5-5 | claude-sonnet-5-5 |
| Input / output | $4 / $20 per MTok | $2 / $10 per MTok |
| Cache read | $0.20 | $0.20 |
| Cache write | $5 / $8 (5 min / 1 h) | $2.50 / $4 |
| Claude Code default | Yes | /model sonnet (medium) |
| API default effort | medium (Opus family) | high |
| Context / max out | 1M / 128k class | 1M / 128k (300k batch beta) |
Cache reads are the same $0.20. Agent loops that are cache-hit heavy still cut the Opus premium versus a naive 2× sticker — walk through that math in what a Claude Code task costs on Opus 5.5. Sonnet cache writes are cheaper, so cold starts favor Sonnet more than hot loops.
Prompting still differs: Opus 5.5 wants whole-task handoff (prompting guide). Sonnet 5.5 wants scope, acceptance checks, and between_tools — the five breaking API changes (thinking: disabled is gone, no forced tool_choice any/tool, computer-use toolset IDs).
Eval your own ladder with build-eval / hillclimb. If train moves and test does not, you overfit Anthropic’s Terminal-Bench story.
How to pick this week
Stay on Opus 5.5 when:
- The job is Claude Code default work: ambiguous, multi-PR, security-sensitive, long-horizon.
- You were about to set Sonnet to max “to match Opus.” Pay Opus Xhigh instead if you need AA 56 on their mix.
- CursorBench / FrontierCode / HLE / OSWorld / Chartography / hard review is the grade.
- You already tuned Opus prompting and cache.
Switch explicit /model sonnet (medium) or API Sonnet high when:
- Bugs, features, docs, slides, spreadsheets with a written spec.
- High-volume agents with a spend cap (AWS’s framing).
- You will not enable Max as a culture.
- You are comparing against GPT-6 Astra on coding volume, not against Opus.
Do not “average” Anthropic 70.6%, AA 64%, and Vals 53% into one Terminal-Bench number. Pick one harness and one effort, then read benchmarks as products.
Related reading
- Claude Sonnet 5.5 building guide
- Claude Opus 5.5 launch benchmarks
- Sonnet 5.5 vs GPT-6 Astra
- Opus 5.5 task cost, cache, effort
- Opus 5.5 prompting
- Claude Code build-eval and hillclimb
- Fable 5.1 vs Opus 5.5
- Official: Building with Sonnet 5.5 · AA on Sonnet 5.5 · Sonnet 5.5 on Amazon Bedrock
Vendor table from Anthropic, September 28, 2026. Intelligence Index and $/task from Artificial Analysis model cards and the Sonnet 5.5 article as of September 29, 2026. Vals and CodeRabbit figures are separate harnesses — not comparable to Anthropic percentages.
