Epoch AI's Capabilities Index now lists Claude Opus 5.5 at 167 — first among 253 tracked models. The exact figure on Epoch's page is 167.35, which puts it 0.84 points ahead of OpenAI's GPT-6 Astra. Opus 5.5 is a closed-weights model released on September 22, 2026.
Rank-one headlines travel fast, and this one deserves a closer look than "Anthropic wins." The lead is small, the index is a composite, and the model page itself shows Astra ahead on one of the hardest math benchmarks. Here is what the score actually tells you and how to use it.
TL;DR — what people are asking
| Question | Answer |
|---|---|
| What is the score? | ECI 167 (167.35 on Epoch's page), ranked 1 of 253 |
| Who is second? | GPT-6 Astra, 0.84 points behind |
| When was Opus 5.5 released? | September 22, 2026 |
| Weights? | Closed |
| API price? | $4 per 1M input tokens, $20 per 1M output tokens (per Epoch) |
| Where does Astra win? | FrontierMath Tier 4: 98% vs 95% |
| Does a #1 rank mean pick it? | It means shortlist it; test on your own workload |
| Where do open models sit? | Kimi K3 is listed 13th with lower scores across domains |
What the Epoch Capabilities Index measures
ECI is Epoch AI's attempt to put many benchmarks on one scale. It combines results from more than 50 evaluations into a single number so that models released at different times, and tested on different mixes of benchmarks, can be compared without eyeballing a dozen leaderboards.
Two properties of a composite like this are worth keeping in mind:
- It is relative. The number has no natural unit. Epoch says the useful quantity is how models move against each other, not the raw value.
- It is fitted. Because it blends many benchmarks, its construction involves modelling choices. A re-fit, a new benchmark, or one saturated test dropping out can shift rankings by a point or two without any model changing.
That second point is why a 0.84-point lead should be read as "tied at the top, with Opus 5.5 narrowly ahead today," not as a decisive victory. For a deeper primer on how to read leaderboards without being fooled, see our guide to how to read AI benchmarks and the broader AI benchmarks complete guide.
What Epoch's model page lists for Opus 5.5
The model page gives a short, checkable scorecard:
| Benchmark | Opus 5.5 |
|---|---|
| ECI | 167 (rank 1 of 253) |
| FrontierMath, Tiers 1–3 | 91% |
| FrontierMath, Tier 4 | 95% |
| GPQA Diamond (science) | 91% |
| MirrorCode (software engineering) | 77% |
| SimpleQA Verified (world knowledge) | 72% |
| Earthborne Rangers | 71% |
Epoch lists training compute, parameter count and knowledge cutoff as unknown, which is normal for closed models.
Against GPT-6 Astra, Epoch's page describes Opus 5.5 as matching on several benchmarks, with Astra ahead on FrontierMath Tier 4 (98% against 95%). In other words the composite gap comes from many small differences, not from one dominant result.
Opus 5.5 versus GPT-6 Astra: how to think about the gap
If you are choosing between the two, the ECI lead is the least useful number on the page. What matters is which model does your job at the lowest total cost.
- Math-heavy and proof-style work. Astra's FrontierMath Tier 4 lead is the clearest signal in Epoch's own data. If your work lives at that frontier, test Astra first.
- Software engineering. Opus 5.5 scores 77% on MirrorCode. The explainx.ai coverage of how the model behaves in coding agents — including task cost, caching and effort settings in Claude Code — is a better guide than any single benchmark.
- Cost. At $4 input and $20 output per million tokens, Opus 5.5 is priced as a top-tier closed model. Compare per completed task, not per token; a model that finishes in fewer turns can be cheaper at a higher rate.
- Everything else. For general work the two are close enough that tooling, latency, and your existing integrations will decide it.
We compare the two directly in GPT-6 Sol versus Claude Opus 5.5 and cover Astra on its own in the GPT-6 Astra launch guide. If you are weighing Anthropic's other models, Fable 5.1 versus Opus 5.5 and Opus 5.5 versus Sonnet 5.5 cover the internal choices.
Where the other leaderboards disagree
Epoch's index is not the only scoreboard, and they do not always agree. A benchmark tracker that mirrors the Artificial Analysis Intelligence Index listed GPT-5.6 Sol at the top of that index in October 2026, at 58.9 percent. That is a different index, built from a different benchmark set, and I have not verified that figure against Artificial Analysis directly, so treat it as an illustration of divergence rather than a data point to cite.
The practical lesson is the same one the benchmarks guide makes: a model can lead one composite and trail another. When two credible indices disagree about first place, the models are close, and your workload decides.
Where open weights sit
Epoch's page lists Moonshot's Kimi K3 at 13th with notably lower scores across domains. That is a useful calibration for anyone considering open weights as a drop-in. The open models can be excellent on cost and control — we covered the gap in Mozilla's open-weight frontier gap analysis and the Kimi K3 open weights release — but the top of an aggregate index is still closed.
How composite indices mislead
A single number is convenient, and it hides things you need to know. Four failure modes are worth keeping in mind whenever you read a composite score like ECI.
- Averaging across unlike tasks. A model that is exceptional at math and merely good at coding can tie a model that is strong at both. The composite cannot tell you which you have. If your work is coding, read the coding benchmarks, not the blend.
- Saturation. When many models score near the ceiling on a benchmark, that benchmark stops separating them. Index builders handle this in different ways, and the choice shifts rankings at the top.
- Contamination and tuning. Public benchmarks leak into training data and shape what labs optimize. Scores on hard, newer tests such as FrontierMath Tier 4 are generally more informative than scores on older, widely trained-on tests.
- Cost blindness. An index measures capability, not price or latency. A model one point ahead that costs twice as much per task is not the better choice for most jobs.
A good habit is to record, for each decision you make from a leaderboard, what you assumed and what you later measured. After a few cycles you learn how much weight your own workloads give to published numbers.
What to do this week
- Shortlist, do not switch. Add Opus 5.5 to your evaluation set if it is not there already.
- Run your own tasks. Pick 20–50 real prompts or tickets from your workload and score outputs blind. Even a small blind test beats a leaderboard.
- Measure cost per accepted result. Include retries, tool calls and caching. Our Opus 5.5 task-cost guide shows how to count it.
- Re-check in a month. Composite indices move as benchmarks are added. A lead this thin can flip with the next release.
- Watch the efficiency angle. OpenAI's response to cost pressure is covered in GPT-6.1 Sol and the Astra cost cut.
What people are asking
Is Opus 5.5 "the best model in the world" now?
On this one composite, it is first. Calling it best overall overstates a 0.84-point margin. Different tasks, different tooling and different prices will put different models on top for different teams.
Why does a close lead still make headlines?
Because a composite rank is easy to quote. The honest summary is "tied for the lead," and that is less shareable than "number one."
Does the ECI include agentic or tool-use performance?
It combines more than 50 benchmarks, and Epoch's page shows software-engineering and knowledge tasks among them. I did not find a breakdown in the sources reviewed that isolates agentic tool-use performance, so do not assume the index predicts how a model behaves in your agent harness.
Will the score change?
Possibly. Indices like this are updated as new benchmarks are added or as scores are re-fit. Expect small shifts.
Where can I check the number myself?
Epoch AI publishes the model page for Opus 5.5 and a benchmarks hub. Read the live page before quoting a figure, because the index can be updated.
Honest limitations
- I read Epoch's model page and secondary coverage; I did not reproduce any benchmark.
- Training compute, parameter count and cutoff are not public, so no claim here depends on them.
- The Artificial Analysis comparison is from a third-party tracker and is unverified.
- A one-point composite lead is not a substitute for a workload test.
Related on explainx.ai
- Claude Opus 5.5 launch: benchmarks and pricing
- GPT-6 Astra launch: benchmarks and pricing
- Fable 5.1 vs Opus 5.5
- Opus 5.5 vs Sonnet 5.5
- GPT-6 Sol vs Claude Opus 5.5
- How to read AI benchmarks
- AI benchmarks complete guide
- Opus 5.5 task cost in Claude Code
Primary source: Epoch AI — Claude Opus 5.5
Scores and rankings reflect Epoch AI's model page as of October 3, 2026 and may be revised.
