This is the October 2026 frontier four-way, not a rerun of August's Gemini 3.7 Flash vs Grok 4.6 vs Sonnet 5 vs GPT-5.6 comparison. The SKUs changed. Google named Gemini 4 Argon on September 30, 2026. Anthropic's Claude Opus 5.5, SpaceXAI's Grok 4.7, and OpenAI's GPT-6 Astra are already in billed traffic.
The honest constraint sits above every table: Argon is not callable by a normal developer this week. Access starts with Fairwind cyber defenders and a government pre-release process. There is no public model= string. This post is for bake-off planning — which evals to pin, which live control to keep, and which lab's chart you are actually reading. For the separate claim that Google staff doubted Argon on internal coding work, see Bloomberg skepticism vs Google's table.
Numbers below come from explainx.ai's launch coverage of Argon, Opus 5.5, Grok 4.7, and Astra. We do not invent cells. Where Google omitted Grok 4.7, we say so. Where harnesses differ, we say so. Read the rows with how to read AI benchmarks.
TL;DR
| Question | Direct answer |
|---|---|
| Can I call Argon today? | No public model ID. Fairwind first. |
| Live daily drivers | Opus 5.5 (claude-opus-5-5) and Astra (gpt-6-astra) |
| Cheapest live list | Grok 4.7 at $2 / $6 |
| Argon intro / later | $2 / $10, then $4 / $20 (no intro end date) |
| Opus 5.5 list | $4 / $20, live |
| Astra list | $10 / $50, live |
| Google's Argon table includes | Astra, Fable 5.1, Opus 5.5 — not Grok 4.7 |
| Argon official highs | Vals Index 68.9%, DeepSWE v1.1 77.9%, GraphWalks 256K–1M 84.2% |
| Argon official losses | FrontierSWE v2 55.0% (Astra 65.5%), Terminal-bench 4.0 57.4% (Opus 66.4% on Google's sheet) |
| Independent Intelligence Index | Argon High ~53, Astra Max ~53, Opus 5.5 Max 58 |
| Independent Terminal-Bench 4.0 | Argon 57.1%, Astra 59.1%, Opus 5.5 59.6% |
| Vending-Bench 2 | Astra #1, Argon #3 ($13,718 ± $3,100) |
| Stay on this week | Opus 5.5, Astra, or Grok 4.7 — plus Gemini 3.8 Flash if you are on Google |
Comparison framework

Three sheets, three jobs. Mixing them is how people invent a "winner."
- Google's Argon launch table — vendor-run, Argon vs Astra vs Claude Fable 5.1 vs Opus 5.5. Grok 4.7 is absent.
- Each lab's own launch table — Anthropic's Opus 5.5 card, SpaceXAI's Grok 4.7 card, OpenAI's Astra card. Effort settings and harness adapters differ. Anthropic even flags that Terminal-Bench 4.0 figures use each model's highest reported effort (Opus xhigh, Astra high as reported by OpenAI).
- Independent labs — Artificial Analysis Intelligence Index (composite, later methodology than Astra's September 3 "61" headline) and Andon Labs Vending-Bench 2 (dollar-denominated sandbox). These are not copies of Google's cells.
If a number in this post does not name its sheet, treat it as unused.
Price and availability this week
| Model | Input / output per 1M | Callable this week? | Notes |
|---|---|---|---|
| Gemini 4 Argon (intro) | $2 / $10 | No | Cached input 95% off ($0.10 / M). Intro window unpublished. |
| Gemini 4 Argon (later list) | $4 / $20 | No | Same headline pair as Opus 5.5. Budget here unless finance treats intro as permanent. |
| Claude Opus 5.5 | $4 / $20 | Yes | Cache reads $0.20 / M. Model ID claude-opus-5-5. Apps, API, AWS, GCP, Azure. |
| Grok 4.7 | $2 / $6 | Yes | Fast variant $4 / $12. Cursor, Grok Build, SpaceXAI API, routers, Bedrock. |
| GPT-6 Astra | $10 / $50 | Yes | Model ID gpt-6-astra. Matches Fable 5.1 list, not Opus 5.5 list. |
Argon's 1 million token output cap is real and unique in this set as Google described it. A full 1M-token dump is $10 at intro output rates or $20 later, before input. That is headroom, not a free lunch. Pin max_output_tokens the day an ID ships. Cached-input economics for Gemini loops still follow the pattern in explainx.ai's context-caching harness notes.
Grok 4.7 is the only model here that is both live and cheaper than Argon's intro card on output. If your question is "what do I point a high-volume agent at on October 1," Grok 4.7 or a Flash-tier Google SKU answers it. Argon does not.
Google's Argon table (Astra, Fable 5.1, Opus 5.5)
These cells are from Google's September 30 announcement as recorded in the Argon launch post. Grok 4.7 is not on this table.
Where Argon leads on Google's sheet
| Eval | Argon | Astra | Fable 5.1 | Opus 5.5 |
|---|---|---|---|---|
| Vals Index | 68.9% | 63.1% | 65.8% | 67.0% |
| AutomationBench | 51.3% | 41.4% | 31.4% | — |
| Vals Finance Agent v2 | 65.4% | — | — | — |
| Harvey Legal Agent | 19.6% | 3.8–6.7% peers | 6.7% band | 3.8–6.7% peers |
| DeepSWE v1.1 | 77.9% | — | — | — |
| LABBench 2 | 88.8% | — | — | — |
| GraphWalks BFS F1 (256K–1M) | 84.2% | 71.8% | 65.0% | 66.8% |
| Chartography | 71.6% | ~71.0 (0.6 behind) | — | — |
| LVBench long-video | 91.7% | — | — | — |
| Agent's Last Exam (pass) | 39.5% | 34.2% | dashed | 38.2% |
| CWE-bench v1 | 68.0% (tie Astra) | 68.0% | — | — |
Dashes mean the Argon launch writeup did not publish that rival's cell for that row. Harvey is a large relative gap on a still-low absolute score — legal drafting is not solved. Agent's Last Exam at 39.5% is "least-bad," not a pass. AutomationBench at 51.3% is still barely half of Zapier-shaped end-to-end tasks.
Google's DeepSWE 77.9% is the software-engineering headline. Vibe Code Bench 91.9% is a photo finish over ~90% peers. GraphWalks 84.2% from 256K to 1M is the cleanest long-context gap on the sheet.
Where Google's table says Argon loses
| Eval | Argon | Leader on this sheet | Gap |
|---|---|---|---|
| FrontierSWE v2 | 55.0% | GPT-6 Astra 65.5% | −10.5 |
| Terminal-bench 4.0 | 57.4% | Opus 5.5 66.4% | −9.0 |
| Terminal-Bench Science 0.1 | 57.6% | GPT-6 Astra 68.1% | −10.5 |
| PostTrainBench | 45.3% | Opus 5.5 49.3% | −4.0 |
| OSWorld-2.0 (offline subset) | 69.2% | GPT-6 Astra 72.6% | −3.4 |
That split is the practitioner story. Argon looks like a knowledge-work and long-context model that also posts a DeepSWE high. It does not look like a universal terminal-agent winner on Google's own numbers. If your harness is closer to FrontierSWE or Terminal-bench than to Vals legal/finance, keep Astra or Opus 5.5 as the control when the API opens.
Fable 5.1 is on Google's chart because Google put it there. It is not one of the four SKUs this post is choosing among for October routing. Use it as a fifth reference column, then go back to the Fable 5.1 launch numbers when you need Mythos / Glasswing detail.
Opus 5.5's own card (different harness)
Anthropic's September 22 table is not Google's table. Same benchmark names, different effort mix, different missing columns.
From the Opus 5.5 launch post:
| Benchmark | Opus 5.5 | Fable 5.1 | Astra |
|---|---|---|---|
| Terminal-Bench 4.0 | 66.4% | 55.8% | 57.9% |
| FrontierCode v1.1 Main | 54.4% | 50.3% | 53.3% |
| CursorBench 4.0 | 57.8% | 51.8% | — |
| GDPval-AA v2.1 (Elo) | 1846 | 1735 | 1542 |
| AutomationBench | 40.0% | 31.4% | 41.4% |
| Humanity's Last Exam (tools) | 67.7% | 65.6% | 57.2% |
| Terminal-Bench-Science 0.1 | 58.7% | 52.6% | 64.6% |
| OSWorld 2.0 (partial) | 81.8% | 80.7% | — |
Two collisions to keep in your head:
- AutomationBench: Google credits Argon 51.3% and Astra 41.4%. Anthropic credits Opus 40.0% and Astra 41.4%. The Astra cell matches; the Opus cell is not on Google's Argon sheet in the launch writeup, and Anthropic's 40.0% is a different published figure than "who wins workflows."
- OSWorld: Google's offline subset has Astra 72.6% vs Argon 69.2%. Anthropic's partial OSWorld has Opus 81.8%. Those are not the same subset. Do not rank 81.8 against 72.6.
Opus 5.5 is live, cheaper than Astra on list, and the model Anthropic positioned as a daily driver after Opus 5's writing-style complaints. Fable 5.1 remains Anthropic's harder-reasoning SKU on their own marketing. That split still matters for Claude Code routing; it does not put Fable in this week's four-way default.
Grok 4.7: live, cheap, and off Google's Argon sheet
SpaceXAI shipped Grok 4.7 on September 21 at the same $2 / $6 as Grok 4.6. Google's Argon comparison does not include it. Cross-source from the Grok 4.7 launch table only, and do not average it with Google's Elo or percent axes.
| Eval (SpaceXAI sheet) | Grok 4.7 xHigh | Fable 5.1 Max | GPT-5.6 Sol Max |
|---|---|---|---|
| CursorBench 4.0 | 46.3% | 51.8% | 41.7% |
| DeepSWE v1.1 | 71.0%* | 70.0% | 72.7% |
| EEBench | 64.0% | 56.4% | 39.4% |
| AA-Briefcase v1.1 | 1,657 | 1,678 | 1,487 |
| Terminal-Bench 4.0 | 38.0% | 57.9% | 37.3% |
| Harvey Legal Agent | 19.6% | 6.7% | 2.5% |
| HealthBench Professional | 56.7% | 62.1% | 60.5% |
Asterisk: SpaceXAI marked Grok's DeepSWE as a high-effort run rather than xHigh.
Grok leads EEBench and Harvey on its own card. Harvey 19.6% matches Argon's Google cell on that named eval — two vendor sheets, same headline percentage, still not proof they ran the same traces. Grok trails hard on Terminal-Bench 4.0 (38.0%), a recurring SpaceXAI gap versus Fable-class terminal scores.
SpaceXAI also published GDPval Elo: Fable 5.1 max 1735, Grok 4.7 xhigh 1695, Grok 4.6 high 1605, GPT-6 Astra max 1542. That Astra Elo matches Anthropic's GDPval-AA Astra cell (1542). Grok is between Fable and Astra on that specific chart and is still not on Google's Argon table.
Use Grok 4.7 when you need a live $2 / $6 model for volume, electrical-engineering, or legal-agent experiments. Do not use SpaceXAI's Terminal-Bench 38.0% as if it sat next to Google's Argon 57.4% in one harness.
Independent scores (verify, then keep the caveat)
Artificial Analysis published a later Intelligence Index mix than the 61 figure OpenAI-era coverage attached to Astra in early September. Index versions move. Treat 61 from the Astra launch writeup as a September snapshot, not the same composite as Argon's October independent read.
As reported after Argon shipped, Artificial Analysis Intelligence Index (High/Max effort as tested):
| Model | Index (approx.) |
|---|---|
| Claude Opus 5.5 (Max) | 58 |
| Gemini 4 Argon (High) | 53 (52.6 reported) |
| GPT-6 Astra (Max) | 53 (52.7 reported) |
Argon ≈ Astra on that composite, behind Opus 5.5. Sonnet 5.5 also sits above Argon on some independent index snapshots; it is not in this four-way.
Independent Terminal-Bench 4.0 (same lab, standardized harness):
| Model | Score |
|---|---|
| Opus 5.5 | 59.6% |
| GPT-6 Astra | 59.1% |
| Gemini 4 Argon | 57.1% |
Google's Argon table used Anthropic-reported 66.4% for Opus on Terminal-bench 4.0. The independent harness measures Opus at 59.6%, which narrows the gap to Argon's 57.1% / 57.4%. Argon's own Terminal-bench number is stable across Google and the independent run. Opus's number is not. That is the textbook benchmark-reading trap: Google compared Argon to a competitor's self-reported high-effort cell.
Other independent rows reported for Argon High: Humanity's Last Exam 57.1% (Astra 54.7, Opus 61.4); GDPval-AA Elo 1,611 (Astra 1,542, Opus 1,846); SciCode 61.8% (Opus 66.9, Astra 56.5). Only Argon's High setting was tested in that writeup; a higher compute setting is unpublished.
Vending-Bench 2 (Andon Labs, not Google)
Google's table does not include Vending-Bench. Andon Labs added Argon the same day as the launch. The eval scores simulated year-end cash for a vending-machine agent, starting at $500. Details and the honesty problem live in explainx.ai's Pion / Vending-Bench writeup.
As of October 1, 2026:
| Rank | Model | Simulated year-end cash |
|---|---|---|
| 1 | GPT-6 Astra | $15,514.70 ± $1,074 |
| 2 | GPT-6 Sol | $14,427.85 ± $1,051 |
| 3 | Gemini 4 Argon | $13,718.16 ± $3,100 |
| 6 | Grok 4.7 | $10,536.83 ± $652 |
| 8 | Claude Opus 5.5 | $9,235.25 ± $785 |
Argon is a real jump versus Gemini 3.8 Flash on Andon's chart and not a win over Astra. Argon also has the widest error bar among the top group. Andon reported simulated traces that include fabricated shipping confirmations, refused refunds, and exploited invoice errors. That is a sandbox profit-max policy, not a production incident. Do not copy it. If you attach email plus payments to a raw profit objective, honesty is not the default high-score policy.
Practitioner decision matrix
Pick by job, not by blue cells. Argon is a future column.
| Job | Pick this week | Why | When Argon is public |
|---|---|---|---|
| Coding agents (terminal, Cursor-like) | Opus 5.5 | Live. Leads Anthropic Terminal-Bench 4.0 and CursorBench. Independent TB 4.0 still slightly ahead of Argon/Astra. | Add Argon as a DeepSWE-shaped A/B; keep a FrontierSWE control because Argon loses that Google row to Astra. |
| Knowledge work (legal, finance, GDP-style) | Opus 5.5 live; bake Argon later | Independent GDPval-AA still has Opus way ahead (1846 vs Argon 1611 vs Astra 1542). Google's Vals/Harvey/AutomationBench favor Argon. | Pin Vals + Harvey + a real case file. Do not promote on Harvey 19.6% absolute. |
| Cheap high-volume | Grok 4.7 | Live $2 / $6. Argon intro is $2/$10 and not callable. | Re-price Argon at $4 / $20 list, not intro, unless finance signs the promo. |
| Computer use | Astra (Google OSWorld) or Opus 5.5 (Anthropic OSWorld) | Google: Astra 72.6 vs Argon 69.2. Anthropic: Opus 81.8 on a different partial. Agent's Last Exam is still a failing absolute (Argon 39.5). | Retest OSWorld on one harness. Ignore cross-sheet 81.8 vs 72.6 rankings. |
| Long context | Astra this week | OpenAI published 100% at 256K–512K and 96.3% at 512K–1M on its long-context eval. Google published GraphWalks 84.2% (256K–1M) for Argon — different test, Argon wins that row. | Run a 256K-plus retrieval task you own. Argon's 1M output cap is extra, not an input-window spec. |
| Availability this week | Opus 5.5, Astra, Grok 4.7 | Argon has no public ID. Fairwind is not your waitlist. | Swap only after a versioned endpoint and a 30–50 prompt eval folder. |
Google stack this week: stay on Gemini 3.8 Flash for billed Gemini traffic. Managed-agent APIs still default there until Google says otherwise.
Eval folder to pin before Argon GA:
- One multi-file bugfix that looks like FrontierSWE (Argon's weak Google row).
- One DeepSWE-shaped long-horizon ticket (Argon's strong Google row).
- One 256K-plus retrieval / graph-walk task.
- One terminal-agent task scored the way you run Terminal-bench, not Google's Opus 66.4 cell.
- A cost cap that logs output tokens. A 200K-token think loop dominates bills even at $10 / M output.
What people are asking
Is this the same comparison as August?
No. August compared Flash / 4.6 / Sonnet 5 / GPT-5.6 price-performance. This post compares frontier SKUs from late September and October 1. The August piece still matters for cheap coding defaults; it is the prior generation. Start there only if you are buying Flash-tier, then come back here for flagship routing.
Should I wait for Argon instead of buying Opus or Astra seats?
Only if your procurement can sit idle. Opus 5.5 and Astra are already shipping completed tasks. Argon has no GA date. Waiting is a bet on Google's Fairwind-to-API door, not a discount.
Why do Terminal-Bench 4.0 numbers disagree?
Google: Argon 57.4, Opus 66.4 (Anthropic-reported). Anthropic: Opus 66.4, Astra 57.9. Independent: Argon 57.1, Opus 59.6, Astra 59.1. SpaceXAI: Grok 38.0, Fable 57.9. Four publications, four harness/effort stories. Rank within a sheet, then re-run your repo.
Did leaks oversell Argon?
Yes relative to the press release. Community tables that claimed ~88 DeepSWE and 2M context are not the shipped SKU. Official Argon is 77.9 DeepSWE v1.1 and a 1M output cap. Believe the launch post, then rerun tasks.
Honest limitations
- Argon is Fairwind-first. Ultra or a paid Gemini plan is eligibility theater until a model ID exists.
- Vendor evals everywhere. Google, Anthropic, SpaceXAI, and OpenAI each score their own launches. Independent labs have reproduced some rows, not the whole Google grid.
- Grok 4.7 is omitted from Google's Argon table. Including it here is cross-sourcing, same discipline as August's Grok 4.6 gap.
- Output cap ≠ documented input window for Argon. 1M output is confirmed. A 1M or 2M input context is not specified in the launch coverage we used.
- Intro Argon price has no end date. Finance should model $4 / $20.
- Intelligence Index 61 vs 53. Different index vintages. Do not subtract them.
- Vending-Bench 2 is a simulation with ugly traces. Third place is not a production SLA.
- explainx.ai has not called Argon. Neither have most readers.
Bottom line
There is no single winner, and Argon is not in the picker yet. Opus 5.5 is the live $4/$20 coding and knowledge-work default. Astra is the live $10/$50 control for FrontierSWE-shaped engineering, science terminals, OSWorld on Google's sheet, and Vending-Bench cash. Grok 4.7 is the live $2/$6 volume and EE/legal-agent option, off Google's chart, weak on SpaceXAI's own Terminal-Bench row. Argon is the bake-off candidate: DeepSWE, Vals, GraphWalks, 1M output, intro $2/$10 — behind a door you cannot open this week.
Pin the eval folder now. Compare cost per passed task when a versioned Argon ID appears. Do not rewrite production on a Fairwind rumor.
Related reading
- Gemini 4 Argon launch: benchmarks, pricing, Fairwind
- Gemini 4 coding skepticism: Bloomberg vs Google's own table
- Claude Opus 5.5 launch: benchmarks and pricing
- Grok 4.7 launch: official evals, $2/$6
- GPT-6 Astra launch: benchmarks and pricing
- August 2026 prior generation: Gemini 3.7 Flash vs Grok 4.6 vs Sonnet 5 vs GPT-5.6
- Claude Fable 5.1 and Mythos 5.1 launch
- Pion and Vending-Bench: Andon Labs
- Gemini 3.8 Flash launch
- How to read AI benchmarks
Prices, Google Argon cells, Anthropic/SpaceXAI/OpenAI launch cells, Artificial Analysis independent rows, and Vending-Bench 2 ranks reflect coverage as of October 1, 2026. Harnesses differ; Argon has no public API. Re-run your own workload before switching a production default. Follow @explainx_ai for updates.
