Opus 5 Hits 30.2% on ARC-AGI-3 — What the Jump Means
Claude Opus 5 scores 30.2% on ARC-AGI-3 — ~4× GPT-5.6 Sol Max and ~20× Opus 4.8. ARC Prize verified results, cost/task, Fable’s absence, and the HN contamination debate.
The day Claude Opus 5 launched, ARC Prize published a number that swallowed the rest of the release narrative: 30.2% on ARC-AGI-3 at High effort — verified independently, not just Anthropic’s launch chart.
On the public ARC-AGI leaderboard, that is not a small bump. The previous frontier cluster sat near GPT-5.6 Sol (Max) ~7.8%, with Opus 4.8 (High) at 1.5% and most other systems under 1%. Hacker News noticed — the ARC-AGI Leaderboard thread hit ~120 points and immediately split into breakthrough vs benchmaxxing.
This explainx.ai post is the builder read: what the table says, what ARC-AGI-3 actually tests, cost, why Fable is missing, and how to treat the HN contamination fight without cope or worship.
ARC’s writeup: Opus 5 (High) is the highest-performing model on ARC-AGI-3 as of July 24, 2026; it completed five additional Public Demo environments no model had previously beaten. Max was competitive on ARC-AGI-1/2 but ARC-AGI-3 was only evaluated at High because of the short pre-launch testing window.
That last caveat matters for fairness: Sol’s 7.8% is a Max ladder point; Opus 5’s 30.2% is High. The gap is still enormous — but it is not an effort-matched duel.
Anthropic’s own launch materials already highlighted the same ARC-AGI-3 step-change on a cost/score plot — see our Opus 5 launch charts.
What ARC-AGI-3 Is (and Isn’t)
ARC Prize’s framing: ARC-AGI evolved from passive fluid intelligence (1 and 2) to interactive adaptation (3). Agents face novel environments; they must explore, form hypotheses, and act efficiently — often scored relative to human median steps.
Important protocol notes that show up in both official docs and the HN thread:
Primary leaderboard rows are model + CoT / reasoning settings, not full Claude Code / Codex-style harnesses.
Community / Kaggle tracks exist under different constraints (e.g. compute budgets).
Semi-private evaluation depends on retention assurances so tasks aren’t quietly absorbed into future training.
So when commenters say “test Claude Code, not Opus,” they are asking for a different benchmark class — agentic coding suites like Frontier-Bench, not a replacement for ARC’s chosen protocol.
Why the Jump Feels Unnatural (and Why That Doesn’t Settle the Argument)
On ARC-AGI-1/2, Opus 5 is strong but in the same league as Sol Max / prior frontiers (high 90s / high 80s–90s). On ARC-AGI-3, it leaves everyone in the dust. That outlier shape is exactly what triggers contamination theories.
Hypothesis A — Real capability on interactive adaptation
ARC Prize and Anthropic-adjacent commentary emphasize stronger logical reasoning → autonomous exploration and planning in unfamiliar envs. Clearing five never-before-beaten public demos is hard to dismiss as a single lucky seed.
HN practitioners who played the public games argue 30% is not “solving the first two easy levels only” in a trivial sense — scoring includes efficiency vs human medians, gotchas on harder levels, and environments many humans fail.
Hypothesis B — Training distribution overlap / “right RL envs”
A common HN guess: the leap is less “general IQ” and more RL / curricula that look like ARC-AGI-3-style interactive puzzles. That can be legitimate capability and still be specialized. Those are not mutually exclusive.
Hypothesis C — Contamination / leakage / naughty system prompts
Threads floated: memorized solutions, stating “hidden rules” then playing byte-identical optimal traces, markdown cheat sheets in system context, session leakage across resets. Some of that discussion mixes public-demo harness claims (e.g. schema-harness 99% talk) with frontier API leaderboard rows — different animals.
ARC’s design (private sets, retention rules, no external harness on the main CoT board) exists specifically to make C harder. It cannot make C impossible for closed models.
explainx.ai posture: celebrate the verified score; refuse both “AGI arrived” and “all benches are fake.” Use private evals for product decisions. Pair ARC with agent harness reality checks.
The Harness Debate (Why Exclusion Exists)
HN split hard:
Camp
Claim
Include harnesses
Real products are systems; pen-and-paper analogy; excluding tools makes the bench less relevant
Exclude harnesses
Otherwise you measure the DSL / search / A* wrapper; “creating the harness is the work”; inductive bias saturates
ARC’s public protocol currently sides with exclusion for the main model board, while allowing models to write tools inside an episode in some setups. That choice keeps the metric closer to “first contact with a novel env” — and farther from “shipped agent product.”
If you care about shipping, track both: ARC-style fluid adaptation and harnessed coding/computer-use benches (model vs effort).
Why Fable 5 Is Missing
Commenters: ARC only runs the semi-private set when providers guarantee retention policies that won’t feed those tasks back into training. Fable’s data-retention story reportedly didn’t clear that bar — so no Fable datapoint, even though Anthropic materials sometimes cite Fable-class public-demo approx. ~20% vs Opus 5’s 30.2% in secondary coverage.
Absence ≠ “Fable can’t do it.” Absence = protocol couldn’t run it under ARC’s trust constraints.
Cost: $20K Feels Like a Third-World SWE — Until You Compare
Leaderboard Cost (V3) for Opus 5 High is ~$20.7K; Sol Max ~$25.1K. Commenters joked that’s a software engineer’s salary in some countries. Fair sticker shock — and also the point of ARC’s cost axes: intelligence without efficiency is incomplete.
Per-task ~$1.45 for Opus 5 High sits near Sol Max’s ~$1.44 while scoring ~4× higher on ARC-AGI-3. That is the efficiency story Anthropic pushed on launch day.
Subscription “included usage” vs API list price muddies casual cost comparisons — another HN rabbit hole. For lab-to-lab fairness, trust ARC’s published cost methodology more than Reddit math.
If your day job is saturated around “Opus 4.5-class CRUD,” you may feel little uplift even when benches move. That is a workload ceiling problem, not necessarily a fake bench.
What Builders Should Do
Don’t route prod solely on ARC-AGI-3 — great signal, narrow skill.
Keep a private interactive suite you never publish (games puzzles, internal games, mystery APIs).
Track effort ladders — High vs Max vs Fast mode economics (developer guide).
Separate model vs harness evals — score Claude Code / Codex sessions distinctly.
Re-read ARC notes when filtering ($10K run caps, preview flags, partial tests).
Assume labs optimize for public boards — that’s rational; your moat is private.
text
Private ARC-style smoke test:
- Novel interactive toy with undocumented rules
- Score: success + steps vs a human baseline
- Never post the env online
- Re-run after each model upgrade
Scores and costs as published by ARC Prize around July 24–25, 2026. Re-verify live leaderboard rows, effort labels, and testing policy before citing in investor or product docs.