explainx.ai0k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsagentsllmsdesignsdictionaryagi trackerranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

On this page

  • TL;DR
  • What actually happened, in order
  • Why this is a process story, not a scandal
  • What this means for model selection
  • What people are asking
  • Related reading on explainx.ai
← Back to blog

explainx / blog

OpenAI Changed GPT-6 Astra's Benchmark Numbers After Launch — Twice

GPT-6 Astra, OpenAI, Benchmarks, AI Trust, Model Selection

Days after launching GPT-6 Astra, OpenAI quietly revised its hallucination rate and a cybersecurity score — then partially reverted one of the changes. Here's exactly what moved, per Fortune's reporting, and why post-launch benchmark edits matter for model selection.

Sep 6, 2026·6 min read·Yash Thakker
add explainx.ai
go deep
OpenAI Changed GPT-6 Astra's Benchmark Numbers After Launch — Twice

OpenAI launched GPT-6 Astra on September 3, 2026 with a published hallucination rate of 4.2%, down from 12.2% for its predecessor, GPT-5.6 Sol. Within days, according to Fortune's reporting, that figure was quietly revised down to 2% — then reverted back toward the original numbers. Separately, OpenAI acknowledged a cybersecurity benchmark comparison used a reasoning tier not commercially available to customers, inflating one of Sol's reported scores. Neither change is dramatic on its own, but together they're a useful, concrete reminder of how provisional a launch-day benchmark table actually is.

This adds real detail to explainx.ai's broader coverage of Astra's rollout and its Critical-tier cybersecurity classification — worth reading alongside this post if you're deciding whether to build on Astra based on its published numbers.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.


TL;DR

table · 2 cols
QuestionShort answer
What changed?Astra's published hallucination rate: 4.2% at launch → cut to 2% → reverted toward 4.2%/12.2%
What else was flagged?A cybersecurity comparison score (11.5% vs. 5.5%) used a reasoning tier not commercially available
TimelineLaunched September 3, 2026; changes reported within the following days
Astra's headline cyber score100% on ExploitBench, versus Sol's 78.5% — not the figure under dispute
SourceFortune's reporting, September 4, 2026
Did it favor Astra?Mostly, per Fortune, though some Anthropic comparison figures moved too
The actionable takeawayLaunch-day benchmark tables are provisional — re-verify before basing a decision on a specific number

What actually happened, in order

  1. September 3, 2026 — OpenAI launches GPT-6 Astra, publishing a hallucination-rate comparison of 4.2% for Astra versus 12.2% for GPT-5.6 Sol, alongside a headline ExploitBench cybersecurity score of 100% for Astra versus 78.5% for Sol.
  2. Within days — OpenAI quietly revises the hallucination figure down to 2%, without a prominent changelog entry accompanying the edit, per Fortune's reporting.
  3. Shortly after — the figure is reverted back toward the original 4.2%/12.2% numbers, and OpenAI separately acknowledges reviewing a cybersecurity comparison score, where an 11.5% figure attributed to Sol reflected a reasoning tier that isn't actually available to paying customers, versus a 5.5% figure under standard, commercially available settings.

Fortune's framing is careful here: the changes "mostly flattered Astra," but the outlet also notes some of Anthropic's own published comparison scores moved during the same reporting window — this isn't a clean story of one company's numbers being static while a competitor's shift.

Why this is a process story, not a scandal

It's worth being precise about what this is and isn't. This isn't documented evidence that OpenAI fabricated a result — it's documented evidence that numbers changed shortly after publication, without the kind of prominent correction notice that would make the revision easy to track unless a reporter happened to catch the before-and-after. That distinction matters. Benchmark methodology genuinely does get revised — a scoring bug found, a reasoning-tier setting clarified, an eval re-run under corrected conditions — and legitimate corrections happen at every lab. The issue Fortune's reporting actually surfaces is transparency about the correction itself, not necessarily bad faith in the original number.

The cybersecurity figure is the sharper example of a real methodological question: comparing a number produced under a reasoning tier that customers can't actually access against a competitor's number produced under standard settings isn't a fair comparison, regardless of intent. That's a genuinely fixable transparency practice — disclose which settings produced which number — separate from whether the underlying capability claim (Astra hitting Critical-tier on OpenAI's own Preparedness Framework, discussed in our earlier coverage) holds up.

What this means for model selection

If you're choosing between frontier models based on a comparison table you saw on launch day, the practical lesson here isn't "distrust OpenAI specifically" — it's "distrust launch-day tables generically, from any lab, until they've had a few days to settle." A few concrete habits this incident argues for:

  • Check the date on any benchmark table you're citing. A screenshot from launch day may already be stale by the time you act on it.
  • Ask which reasoning tier or settings produced a number, especially for cybersecurity, agentic, or reasoning-heavy evals where a "maximum effort" or gated tier can dramatically change a result versus what's actually available in the product you'd be paying for.
  • Independent aggregators — Artificial Analysis, for instance — exist precisely because self-reported launch numbers need a second measurement before they're fully trustworthy for a purchasing decision.

What people are asking

Does this mean GPT-6 Astra's hallucination rate is actually worse than advertised? Not necessarily — the number moved in both directions (down to 2%, then back up), which is more consistent with methodology uncertainty than a one-way inflation designed to make the model look artificially good. Treat the settled figure as somewhere in the originally-reported 4.2% range until OpenAI publishes a clear, dated final number with its methodology attached.

Is Astra's 100% ExploitBench score itself in question? No — that specific figure isn't what's under scrutiny in Fortune's reporting. The disputed number is a comparison figure for the previous model, Sol, used as context alongside Astra's launch, not Astra's own headline result.

Should this affect trust in OpenAI's Preparedness Framework classification specifically? They're separate questions. The Critical-tier cybersecurity classification is a policy and access decision (gating advanced capability behind the Daybreak partner program) documented in OpenAI's own Path to Astra post; the benchmark-number revisions are a separate transparency issue about how comparison figures were presented and corrected. One doesn't necessarily undermine the other, but both are worth tracking independently.

How does this compare to how other labs handle benchmark corrections? Every major lab has revised published numbers post-launch at some point — it's a normal part of the benchmark lifecycle. What makes this instance a story rather than routine housekeeping is that a specific outlet (Fortune) documented the before-and-after with dates and figures close to the event, which is rarer than the underlying practice of quiet revisions itself.


Related reading on explainx.ai

  • Claude Fable 5.1 and Mythos 5.1: Benchmarks, Pricing, and Safeguards — Anthropic's comparable launch-week benchmark disclosure, for contrast
  • OpenAI Confirms Astra Is Critical-Tier for Cybersecurity — Path to Release — background on the Preparedness Framework classification and Daybreak gating
  • GPT-6 Astra Scores 95% on a Robot Control Task — an independently-run benchmark on the same model, for comparison
  • Artificial Analysis Intelligence Index v4.2 — an independent aggregator's read on frontier model rankings
  • AI Benchmarks: The Complete Guide — background on how to read any benchmark table skeptically
  • How to Read an AI Benchmark and Not Get Fooled — a practical checklist directly applicable to this story

Primary source: Fortune — "OpenAI quietly boosts some of Astra's evaluation metrics, and continues to change others post-launch," September 4, 2026 · OpenAI — Path to Astra

This post reflects Fortune's reporting and OpenAI's own Path to Astra documentation as of September 6, 2026. Benchmark figures are subject to further revision by OpenAI — verify current numbers against OpenAI's own published documentation before citing specific percentages in a procurement decision.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

Related posts

Sep 6, 2026

GPT-6 Astra Scores 95% on a Robot Control Task, Up From Fable 5.1's 40%

A Robocurve benchmark thread from Jay Chooi puts GPT-6 Astra well ahead of Claude Fable 5.1 on a robot-arm control task — 95% success versus 40% — while using a fraction of the output tokens. On harder, precision-limited tasks the two models tie, but Astra still gets there cheaper and faster. If the token-efficiency trend holds, LLMs could control robot arms in real time within a year or two.

Sep 5, 2026

Artificial Analysis Intelligence Index v4.2: What Actually Changed

Artificial Analysis published Intelligence Index v4.2 on September 4, 2026, an interim update ahead of v5: two new evaluations added (AA-Briefcase, GDP.pdf), GPQA Diamond retired as saturated, and private held-out test sets now carry 40% of the total weight. Claude Fable 5.1 leads the index, GPT-6 Astra wins on cost-per-task and token efficiency — and Hacker News raised fair questions about the timing.

Sep 5, 2026

GPT-6 Astra vs Claude Fable 5.1: Which Model Wins Where?

OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1 landed days apart at the same API price. Independent scores favor Fable on general intelligence; OpenAI-reported lanes favor Astra on computer use, math, security, and token efficiency. Here's the decision matrix for builders.