explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR — questions people actually ask
  • The Artificial Analysis chart (what the bars actually are)
  • List price vs tokens vs dollars per task
  • Terminal-Bench and the two leaderboards problem
  • Effort ladders: Sonnet climbs, Astra is already high
  • What GPT-6 Astra still owns
  • GPT-6.1 is not a third column
  • How to pick this week
  • Related reading
← Back to blog

explainx / blog

Claude Sonnet 5.5 vs GPT-6 Astra: Who Wins After the AA Chart?

Comparison, Claude, OpenAI, Sonnet 5.5, GPT-6 Astra, Benchmarks

Sonnet 5.5 scores 56 on Artificial Analysis vs GPT-6 Astra’s 53 at max — at $2/$10 vs $10/$50. Token burn and Terminal-Bench decide the real pick.

Sep 29, 2026·8 min read·Yash Thakker
add explainx.ai
go deep
Claude Sonnet 5.5 vs GPT-6 Astra: Who Wins After the AA Chart?

September 29, 2026 — Artificial Analysis put Claude Sonnet 5.5 at 56 on the Intelligence Index (max effort, default fallback). GPT-6 Astra (max) sits at 53, tied with Claude Fable 5.1 (max) in the same bar chart. Opus 5.5 (max) still leads at 58. The chart people are screenshotting is real. The routing decision is not “Sonnet beats the OpenAI flagship, switch everything.”

This is Sonnet 5.5 vs live GPT-6 Astra, not vs cancelled GPT-6.1 Astra. If your Slack thread still says “wait for 6.1,” that SKU is off the October calendar.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR — questions people actually ask

table · 2 cols
QuestionDirect answer
Who is ahead on AA Intelligence Index?Sonnet 5.5 (max + fallback) 56 vs Astra (max) 53. Opus 5.5 is 58.
Who is cheaper on the API sticker?Sonnet 5.5 $2/$10 vs Astra $10/$50. Same sticker as GPT-6 Sol, not the same model.
Who is cheaper per hard index task?Not automatic. AA measured ~$7.60/task for Sonnet 5.5 at max because of output-token volume. Astra uses far fewer tokens per index task.
Terminal-Bench 4.0?AA: Sonnet ~64%, Astra/Opus ~60%. Anthropic table: Sonnet 70.6%. Pick a source.
Default in Claude Code?/model sonnet is Sonnet 5.5 medium from v2.1.284. Default model is still Opus 5.5. See the building guide.
Is 6.1 in this comparison?No.

The Artificial Analysis chart (what the bars actually are)

Artificial Analysis Intelligence Index bars: Claude Opus 5.5 58, Claude Sonnet 5.5 56, Claude Fable 5.1 53, GPT-6 Astra 53, then Opus 5, Muse Spark 1.3, GPT-6 Sol, GPT-5.6 Sol, Grok 4.7

Source: Artificial Analysis public leaderboard screenshot, Intelligence Index (labels truncated on the image: “max with fallback” on the Claude 5.5 rows). Index version in AA’s Sonnet 5.5 model card is v4.3.2 — ten evals including AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1. That is a newer mix than explainx.ai’s v4.2 write-up from Fable-week.

Read the bars as one composite, not as “IQ.” Fable 5.1 and Astra tie at 53 on this snapshot. Sonnet 5.5 is not a new flagship; it is Anthropic’s $2/$10 tier sitting two points under Opus and three above Astra.

AA’s own article is the part the screenshot leaves out: Sonnet 5.5 (max) used about 193k output tokens per Intelligence Index task, the highest they have measured, roughly 7× GPT-6 Astra (max). Quality at that setting is purchased with decode.

List price vs tokens vs dollars per task

table · 3 cols
Claude Sonnet 5.5GPT-6 Astra
RoleEveryday 5.5 tier (launch guide)OpenAI flagship (launch numbers)
Input / output$2 / $10 per MTok$10 / $50 per MTok
vs Sonnet 5Same sticker; Anthropic claims up to ~30% less per task on typical work because of fewer tokens and speedN/A
Context1M native1M class (OpenAI launch; some aggregators show ~1.05M)
AA Intelligence (max)56 (with fallback)53
AA effort ladder (Sonnet)Low 36 → Med 41 → High 47 → Xhigh 52 → Max 56Low 46 → Med 50 → High 51 → Xhigh 52 → Max 53
AA output speed (card)Up to 171 t/s at Xhigh (AA’s fastest Sonnet 5.5 row)About 61 t/s at max (AA card)

Five times cheaper on the sticker is the headline finance teams will quote. It is true for input and output rates. It is false as a promise that a max-effort Sonnet job costs one-fifth of a max-effort Astra job.

AA says Sonnet 5.5 at max costs about $7.60 per Intelligence Index task, ~50% more than Sonnet 5 on that same cost-per-task metric, even though the per-token price did not change. The extra quality is extra tokens. Astra’s AA model table lists a $7.7 “price” column next to 53 intelligence — treat AA’s own article ($7.60/task, 193k output tokens) as the Sonnet cost story, and do not smash unlike columns into one spreadsheet.

Practical billing rule:

  1. High-volume, bounded tasks (bugs, docs, slides, scoped agents) → start Sonnet 5.5 medium/high. That is the job Anthropic staffed the Osmani guide for.
  2. Max effort as a habit → you are paying Opus-adjacent quality with Sonnet’s rate card but Opus-like decode. Measure $/accepted PR, not $/MTok.
  3. Astra still makes sense if your harness is Codex / ChatGPT, you need Astra-specific computer-use history from Astra vs Fable, or AA’s science-terminal split still favors OpenAI on your workload.

For spend math on the Anthropic side, keep Opus 5.5 task cost next to this post. Sonnet cache read is $0.20/M (same as Sonnet 5); Astra cache read was $1.00/M in the Fable comparison — cache-heavy loops tilt Anthropic even before the 5× sticker.

Terminal-Bench and the two leaderboards problem

Anthropic’s Sonnet 5.5 launch table (in our building guide) has Terminal-Bench 4.0 at 70.6% for Sonnet vs 66.4% Opus at Xhigh and no Astra cell.

Artificial Analysis’s Sonnet 5.5 article has Terminal-Bench 4.0 at 64% for Sonnet 5.5 vs 60% for Opus 5.5 and GPT-6 Astra. Same name, different numbers.

Do not average them. Vendor evals and AA evals disagree by six points on Sonnet. Both still put Sonnet ≥ Astra on Terminal-Bench 4.0. That is the only claim that survives both tables.

AA also reports Terminal-Bench-Science (not in the Intelligence Index): Sonnet 53%, behind GPT-6 Astra and Opus 5.5. If your agents look like lab workflows in a shell, the composite index overstates Sonnet vs Astra.

FrontierCode / CursorBench on Anthropic’s chart still favor Opus, and Sonnet drops at Max on FrontierCode because of Claude Code review-skill timeouts. That is a Sonnet vs Opus issue, not Astra — but it is why “max everything” is a bad default even when the AA index loves max.

Effort ladders: Sonnet climbs, Astra is already high

AA’s Sonnet 5.5 family on the Intelligence Index: 36 / 41 / 47 / 52 / 56 from Low → Max. Astra’s family: 46 / 50 / 51 / 52 / 53.

Sonnet Low (36) is not in the same league as Astra Low (46). If you “save money” by slamming Sonnet to Low, you can lose the 56-vs-53 story entirely. Astra’s score is flat across effort compared with Sonnet’s steep climb.

That matches how you should configure APIs:

  • Sonnet: effort is a real knob. Claude Code maps sonnet to medium. API default effort on Sonnet 5.5 is high. Max is an eval setting, not a production default.
  • Astra: you are already near the index ceiling at medium/high. Paying for max buys less index movement than on Sonnet.

What GPT-6 Astra still owns

The Astra vs Fable 5.1 matrix still applies at the flagship price: computer use, some math/security lanes, token efficiency, launch-week GUI demos. Sonnet 5.5 did not delete that product.

Astra vs Sol is the OpenAI-internal trade: Sol is $2/$10, same sticker as Sonnet 5.5, weaker than Astra. Do not swap “Sonnet 5.5 vs Astra” with “Sonnet 5.5 vs Sol.” Sol’s AA bar in the screenshot is 48 (max). Sonnet 5.5 at 56 is a different comparison.

If the question is OpenAI mid-tier vs Anthropic mid-tier, pair this post with Sol vs Opus 5.5 — that is Sol vs flagship Claude, not Sonnet vs Astra.

GPT-6.1 is not a third column

OpenAI shelved October GPT-6.1 Astra after tests on scope authorization, deception, and how the model reports work. Full story: cancellation post. This comparison uses shipping GPT-6 Astra. A 6.1 bar is not on the AA chart you pasted.

How to pick this week

Use Sonnet 5.5 when:

  • You are in Claude Code or the Claude API and the task is scoped (bugs, features, docs, slides).
  • You can live with adaptive thinking and the migration breaks (between_tools, no forced tool_choice, toolset IDs).
  • You will hillclimb on your eval, not AA’s index — build-eval / hillclimb.
  • Sticker $2/$10 matters and you will cap effort so you do not accidentally buy 193k-token max runs.

Use GPT-6 Astra when:

  • The job is already on Codex/ChatGPT and switching vendors is the expensive part.
  • AA Terminal-Bench-Science (or your own science-shell eval) still prefers Astra.
  • You need Astra’s computer-use / efficiency profile from the Fable-week head-to-head.
  • You want a model whose AA score does not depend on lighting money on fire at max.

Use Opus 5.5 when:

  • You were about to set Sonnet to max anyway. At 58 vs 56, Opus is the honest flagship at $4/$20, not a mystery third screenshot bar.

Run one held-out set. If train improves and test does not, you overfit the index. That is the same warning as the hillclimb playbook.

Related reading

  • Opus 5.5 vs Sonnet 5.5
  • Claude Sonnet 5.5 building guide
  • GPT-6 Astra launch benchmarks
  • GPT-6 Astra vs Fable 5.1
  • GPT-6 Astra vs Sol
  • GPT-6.1 Astra October cancellation
  • AA Intelligence Index v4.2 (older mix)
  • How to read AI benchmarks
  • Official: AA on Sonnet 5.5 · Building with Sonnet 5.5

Intelligence Index scores and token counts follow Artificial Analysis’s Sonnet 5.5 article and model cards as of September 29, 2026. Anthropic Terminal-Bench figures follow Anthropic’s launch table. They are different harnesses.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 23, 2026

GPT-6 Sol vs Claude Opus 5.5: Same-Day Launches, Different Price Tiers

OpenAI's GPT-6 Sol and Anthropic's Claude Opus 5.5 launched within hours of each other, and the instinct is to treat them as direct rivals. The pricing tells a different story — Sol is a mid-tier, cost-optimized model at half Opus 5.5's price, not a flagship competing on raw capability. Here's what actually overlaps, and where the comparison breaks down.

Sep 23, 2026

Grok 4.7 vs Claude Opus 5.5 vs GPT-6 Sol: The Only Numbers That Overlap

Three models, three companies, three separate benchmark suites — Grok 4.7, Claude Opus 5.5, and GPT-6 Sol all launched within 48 hours of each other in September 2026, and none of them published a shared eval table against the other two. Terminal-Bench 4.0 is the one benchmark all three companies actually reported, and the gap on it is not close.

Sep 29, 2026

Claude Opus 5.5 vs Sonnet 5.5: Same Family, Different Bill

Anthropic’s launch table puts Sonnet 5.5 within a few points of Opus 5.5 on knowledge-work Elo and ahead on Terminal-Bench 4.0. Artificial Analysis still ranks Opus 58 vs Sonnet 56 — and Sonnet at max costs more per index task. This is the routing matrix for Claude Code and the API.