Update — September 27 staging: Registry leak — Claude Sonnet 5.5 Droid greyscale evidence.
September 28, 2026 — Anthropic’s second Claude 5.5 model is live. @claudeai introduced Claude Sonnet 5.5 (5.6M+ views) as a clear upgrade over Sonnet 5: more than 30% faster, up to about 30% lower cost per task at the same per-token prices, and the everyday counterpart to Opus 5.5. @ClaudeDevs pointed builders to Addy Osmani’s official playbook: Building with Claude Sonnet 5.5.
Official launch videos (@claudeai)
The family intro clip is the 5.6M-view post. The two follow-ups Anthropic pinned in the same thread are the ones builders actually quote in Slack: same list price, fewer tokens, faster output, then design/slides.
Cost and speed (same $2/$10, up to 30% less per task, more than 30% faster):
Design and slides (UI polish, template-following decks):
Developer guide on X (@ClaudeDevs)
| Field | Sonnet 5.5 |
|---|---|
| Model ID | claude-sonnet-5-5 (Bedrock: anthropic.claude-sonnet-5-5) |
| Context | 1M tokens, native |
| Max output | 128k (300k batches with beta header) |
| Knowledge cutoff | June 2026 |
| Thinking | Adaptive on by default; between_tools for tool loops |
| Default effort | high on API; medium in Claude Code |
| Pricing | $2 / $10 per M input/output (unchanged vs Sonnet 5) |
Official evals (Anthropic launch charts)
The digest headline “Sonnet 5.5 beats flagship Opus 5.5 on Terminal-Bench 4.0 at half the price” is half true. List price is $2/$10 vs $4/$20. On Terminal-Bench 4.0, Anthropic’s table really does put Sonnet ahead. On almost every other row, Opus still wins. Read the footnotes before you rewrite a default-model policy.
Figures below are Anthropic’s own launch graphics (Sonnet 5.5 System Card methodology). Attribution: Anthropic, September 28, 2026.

| Eval | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 10.3% | 66.4% (Xhigh) | — |
| FrontierCode 1.1 (Main) | 52.1% Xhigh / 46.2% Max | 42.4% | 54.4% | 49.3% |
| CursorBench 4.0 | 55.5% | 34.1% | 57.8% | — |
| GDPval-AA v2.1 (Elo) | 1844 | 1449 | 1846 | 1487 |
| AA-Briefcase v1.1 (Elo) | 1811 | 1359 | 1822 | 1483 |
| Humanity’s Last Exam (tools) | 64.5% | 54.9% | 67.7% | — |
| OSWorld 2.1 (partial) | 80.1% | 57.0% | 81.8% | — |
| Chartography (no tools) | 61.6% | 15.6% | 64.4% | 53.6% |
Four footnotes that change how you use the table:
- Opus Terminal-Bench is reported at Xhigh, not a mystery “default.”
- FrontierCode Max vs Xhigh: Sonnet drops at Max (46.2% vs 52.1%). FrontierCode grades whether a change could merge without human edits and penalizes out-of-scope diffs even when they are useful. Anthropic says Max more often fires Claude Code’s code-review skill, which fans the review across subagents; Cognition found two cases that timed out or added extra out-of-scope edits. If your harness already runs a review swarm, do not assume Max is better.
- Artificial Analysis ran GDPval-AA and AA-Briefcase on a pre-release Claude Platform deployment that had a structured-output bug. Anthropic says the bug is fixed and any remaining effect understates Sonnet 5.5.
- GPT-6 Sol image-understanding scores may still reflect a Surge AI Chartography / AA snapshot before OpenAI’s later image bugfix. Anthropic does not expect large AA-Briefcase / GDPval moves from that fix; internal Chartography retests looked unimpacted.
Cost vs quality at each effort level
These four plots are the actual @claudeai claim in chart form: Low/Medium Sonnet 5.5 often matches Sonnet 5’s best score at a fraction of the dollars. X-axis is USD per task (log) except Terminal-Bench, which is USD per attempt.

On AA-Briefcase, Sonnet 5.5’s Med → High → Xhigh → Max curve sits on top of Opus until the far-right Max point, where Opus still edges Elo. Sonnet 5 is a lower parallel line. GPT-6 Sol is competitive at cheap Low/Med and then stops on this plot before Claude’s Max.

CursorBench 4.0 does not publish GPT-6 Sol, so Anthropic plots GPT-5.6 Sol. Sonnet 5.5 Low already sits near 36%; High/Xhigh track Opus; Max is still a hair under Opus. Sonnet 5 never leaves the 25–35% band at much higher spend.

FrontierCode is the trap chart. Sonnet 5.5 peaks around Xhigh (~52%) then falls at Max toward ~46%. Opus stays in the mid-50s. If you “just turn Max on” because it sounds like more thinking, this eval punishes you. Pair this with the Claude Code build-eval / hillclimb loop: held-out tasks, one change per round, revert if only the train split moves.

Terminal-Bench 4.0 is where Sonnet Max crosses above Opus. GPT-5.6 Sol is the OpenAI stand-in (no public GPT-6 Sol row). Sonnet 5 is almost flat in the single digits until late spend. This is the chart behind “half the price, higher terminal score” — it is real, and it is one eval.
How to route from these plots: default Sonnet 5.5 medium in Claude Code for scoped bugs and docs; API high for knowledge-work Elo; Xhigh not Max if FrontierCode-style merge-cleanliness is the grade; Opus 5.5 when CursorBench / HLE / OSWorld / Chartography are the job. Re-run your own eval — Anthropic’s harness is not yours.
Sonnet 5.5 vs Opus 5.5 — when to use which
Anthropic’s table (via claude.dev) is the routing doc teams should paste into runbooks:
| Workload | Start with |
|---|---|
| Bug fixes, feature iteration, high-volume dev | Sonnet 5.5 |
| Polished docs, slides, spreadsheets, design-sensitive one-pagers | Sonnet 5.5 |
| Repeatable agent tasks (investigate, review, draft) with clear specs | Sonnet 5.5 |
| Long-horizon agentic coding, hardest judgment calls | Opus 5.5 |
Epic COO Daniel Vogel quoted in the guide: Sonnet 5.5 cleared a system design audit and data-flow review on large gameplay codebases with less prescriptive prompting — early enterprise signal, not a universal benchmark substitute.
For Opus-specific spend math, keep what a Claude Code task costs on Opus 5.5 beside this post; Sonnet 5.5 deserves the same dollars-per-shipped-feature treatment once your cache hit rate is known.
Migration breaking changes (API)
Osmani’s guide lists five breaking changes plus response-shape updates. The ones that break production silently:
1. thinking: disabled → between_tools
# Sonnet 5.5 — tool-heavy agent
client.messages.create(
model="claude-sonnet-5-5",
max_tokens=16000,
thinking={"type": "between_tools"},
output_config={"effort": "high"},
messages=[{"role": "user", "content": "..."}],
)
between_tools only works at low/medium/high effort — not xhigh/max.
2. Forced tool_choice → auto + strict: true
tool_choice types any and tool return 400, including on token counting.
3. Read blocks by type
Default thinking means content[0].text is wrong — loop block.type for thinking vs text.
4. Computer use toolset
Claude API / Google Cloud: use computer_toolset_20260801; legacy computer_20251124 returns 400 (Bedrock still accepts the older declaration per Anthropic docs).
5. Claude Code migration skill
Run /claude-api migrate this project to claude-sonnet-5-5 in Claude Code to apply ID swaps and parameter fixes repo-wide.
Tuning and refusals builders should know
- Re-run effort sweeps — levels are recalibrated vs Sonnet 5; start high on API, medium for agentic coding unless evals say otherwise.
- Drop Sonnet 5 prompt hacks — remove “do not be lazy” shims and re-eval before adding new ones.
- Cyber refusals — Sonnet 5.5 adds cyber safeguards similar to flagship models; declines return
stop_reason: "refusal"with categories likecyberandreasoning_extraction. - Images — high-res tier up to 2576px long edge costs ~2.5× image tokens vs Sonnet 4.x — downscale when detail is unnecessary.
X controversy: reference photo credit
Launch creative included a code-to-painting demo comparing Sonnet 5, Sonnet 5.5, and Opus 5.5 on a window-seat photograph. @IceSolst noted the image matched their September 26 post; @fire argued uncredited commercial use raises liability questions for Anthropic. The guide credits @jkeatn for demo ideas and @IceSolst for the reference — the dispute is about permissions, not model capability. Treat it as a launch comms lesson if you ship model demos with user-generated references.
Claude Code defaults
From v2.1.284:
/model sonnet→ Sonnet 5.5, medium effort, 1M context.- Default model remains Opus 5.5.
- No fast mode on Sonnet 5.5; thinking cannot be fully disabled in Claude Code — effort controls depth.
Pair with Claude Opus 5.5 prompting guide for effort and cache patterns on the same 5.5 family.
What @claudeai said besides the launch clip
The family intro, the cost/speed clip, and the design clip are one thread. The numbers teams actually ship against:
- Role in the family: faster, lower-cost complement to Opus 5.5 — strongest at well-scoped everyday tasks, bug fixes, and polished documents, slides, and spreadsheets.
- Same list price as Sonnet 5, fewer tokens per task — Anthropic’s testing: up to 30% less cost per task, more than 30% faster output. Fastest Sonnet to date.
- Effort economics: on several benchmarks, Sonnet 5.5 at Low or Medium effort beats Sonnet 5’s best score at about a tenth of the cost.
- Design: polish on UIs; follows templates for slides that need minimal editing.
- Writing: like Opus 5.5, clearer than the prior generation; speed favors fast iteration.
- Safety: automated behavioral audit improves or matches Sonnet 5 on most alignment/honesty measures. First Sonnet with cyber safeguards and fallbacks like flagship models. Routine software development unaffected.
- Availability: everywhere today; Haiku 5.5 still “coming weeks.”
That last point is why the Droid leak aged in 24 hours: the slug was real, the marketing switch flipped September 28.
X replies included the expected pacing-the-frontier joke — ship Opus 5.5, then Sonnet 5.5, while asking the industry to slow down. That is a politics thread, not a migration bug. For enterprise routing, ignore the joke and use the workload table above.
Availability IDs (copy into config)
| Surface | ID |
|---|---|
| Claude API / GCP / Foundry / Claude Platform on AWS | claude-sonnet-5-5 |
| Amazon Bedrock | anthropic.claude-sonnet-5-5 |
| Microsoft Foundry | Global Standard deployments only, per Anthropic |
US-only inference (inference_geo: "us") is 1.1× standard price. Cache write $2.50 / $4 (5 min / 1 hour); cache read $0.20 — same as Sonnet 5, half of Opus 5.5 on writes.
Related reading
- Opus 5.5 vs Sonnet 5.5 — AA 58 vs 56, when max Sonnet costs more per task
- Sonnet 5.5 vs GPT-6 Astra
- Claude Sonnet 5.5 registry leak (Sep 27)
- Claude Opus 5.5 launch benchmarks
- Opus 5.5 task cost in Claude Code
- Claude Code build-eval and hillclimb — same-day claude-api skill for evals
- Official: Building with Claude Sonnet 5.5 · @claudeai intro · cost/speed · design · @ClaudeDevs guide
Model IDs, breaking changes, and pricing match Anthropic’s September 28, 2026 claude.dev guide — verify current docs before production cutover.
