Update — September 4, 2026: Sam Altman posted his own apology for the rollout the morning after launch: "first, sorry for the messy rollout. second, when we screw up, we try to make it right. third, we should be able to begin broad rollout to API customers and chatgpt subscribers in the near future. as usual we will start with pro subscribers." That confirms what the "Plus users alongside Pro" framing in the Rollout mechanics section below implied but didn't fully spell out: access on launch day was narrower than the announcement suggested, Pro subscribers are first in the actual queue, and broad API and ChatGPT-subscriber access is still pending as of this update — not simultaneous with the September 3 announcement. In X replies, Altman confirmed no firm ETA beyond "hopeful... this weekend," and the loudest recurring request in the thread wasn't about Astra at all — a sustained push from users asking OpenAI to bring back or open-source the retired GPT-4o model family.
GPT-6 Astra shipped on September 3, 2026 — after its own launch page briefly didn't work.
Multiple outlets, including CNBC, Reuters, The Verge, and VentureBeat, had already published stories quoting OpenAI's official announcement material when regular visitors hit OpenAI's Astra page and got errors. For close to an hour, the only reliably working copy of the announcement was an archived mirror circulating on social media. Then, in the middle of the confusion, Tibo Sottiaux — who leads Codex and ChatGPT at OpenAI — posted the plainest possible confirmation: "We are starting to release GPT-6 Astra and we are doing it as carefully and quickly as possible."
By the next morning, the launch had a name, a price, a benchmark table, and — per Wharton's Ethan Mollick, who had early access — a simulated Library of Alexandria built well enough that people are calling it evidence of real autonomous, multi-day work. Here's everything that's actually confirmed, separated cleanly from the "AGI era" framing already circulating.
TL;DR
| Question | Answer |
|---|---|
| Is it live? | Yes, rolling out — Plus, Pro, Business, Enterprise, plus API/AWS. Full rollout takes "a few days." |
| Price | $10 / $50 per million input/output tokens. Matches Claude Fable 5 and 5.1 exactly. |
| API name | gpt-6-astra |
| Best benchmark | 100% on ExploitBench, 99.2% on SRE-Bench reverse engineering — security is Astra's strongest category. |
| Headline number | 99.9% on ARC-AGI-3 — under OpenAI's own custom harness, disclosed as such. |
| Weakest number | Intelligence Index: 61 — trails Claude Fable 5.1. |
| Long context | 100% accuracy at 256K–512K tokens, 96.3% at 512K–1M. |
| Coding cost | Roughly half of Fable 5's cost per completed coding-agent task. |
| Naming scale | Astra > Sol > Terra > Luna — confirmed by Tibo, bigger celestial object = more capable. |
| Delayed access? | OpenAI is issuing one banked reset per day you're on a paid plan without Astra access. |
What was actually posted
OpenAI's own promotional video opens on a callback: a black-and-white 1980s-style demo of someone asking a computer to draw a yellow circle, and it does — simply, slowly, the way computers actually worked in 1980. Then it cuts to today: OpenAI employees talking to Astra by voice, asking it to turn that same yellow circle into a rocket ship, then into a full playable 3D game, in minutes — and, in one clip, dictating a complete eBay listing from voice alone, no typing. The tagline: "This is GPT-6 Astra. Anything you can do on a computer, Astra can do for you. Fast."
Tibo's own framing, posted at 1:07 AM: "It was very important to us that we bring it to all Plus users and not only Pro, Business and Enterprise... behind the scenes many novel systems will operate at scale for the first time and we are bringing a lot of compute up. It is pure magic." That "bringing a lot of compute up" line is doing real work — it's OpenAI pre-explaining why the rollout is staggered rather than instant, and it's consistent with the sheer scale of infrastructure a frontier launch requires.
He followed up with the naming clarification people had been asking about since GPT-5.6 shipped three tiers in July: "Same as previously — Bigger number = Better, Bigger celestial object = Better — and the scale is Astra > Sol > Terra > Luna." Astra now sits at the top of a scale that started with Sol, Terra, and Luna in July — the naming convention survived the version bump rather than resetting.
The numbers, without the highlight reel
Simon Willison's independent read-through of OpenAI's own published benchmarks is the most useful single source here, precisely because he flags the caveats OpenAI's own marketing wouldn't lead with.
| Benchmark | Astra's score | Note |
|---|---|---|
| ARC-AGI-3 | 99.9% | Under OpenAI's custom "Provider Adapter harness," not the default ARC-AGI harness |
| ExploitBench | 100% | Security/offensive-capability benchmark |
| ExploitGym | 42.4% | Same category, harder task set |
| SRE-Bench (reverse engineering) | 99.2% | |
| Long context, 256K–512K tokens | 100% | |
| Long context, 512K–1M tokens | 96.3% | |
| Intelligence Index | 61 | Trails Claude Fable 5.1 |
| Coding Agent Index | Leads on cost-efficiency | ~half of Fable 5's cost per completed task |
Willison's summary is blunt and fair: Astra is "clearly OpenAI's Fable competitor" with "mixed results — while it excels at benchmarks like ARC-AGI and security tasks, it underperforms Fable on overall intelligence metrics." That's a genuinely useful one-line verdict, and it's worth sitting with rather than skipping past on the way to the AGI framing: this is not an across-the-board win. It's a model that leads hard in two specific lanes — security-adjacent tasks and long-context recall — while trailing the field's other frontier model on the general reasoning metric most people actually care about day to day.
The 99.9% ARC-AGI-3 number deserves the same scrutiny we've applied to every big benchmark claim this year, and here we don't have to guess — the ARC Prize Foundation itself published the breakdown same-day. Under its provider-neutral Standard harness, Astra (max) scores 62.7% for $26,098. Under a Provider Adapter harness — which lets Astra keep its opaque reasoning state and use its own compaction between calls, rather than having that context discarded every turn — the best observed score climbs to 99.9% at high reasoning, for $18,817. ARC's own explanation of the gap is telling: the Standard harness was throwing away Astra's private reasoning after every single action and truncating older moves out of view entirely, forcing the model to "figure out the game anew" repeatedly — a harness limitation, not a capability one. ARC says Provider Adapter runs were also 3.66x faster and used 49% fewer total tokens across the games both harnesses solved. Both numbers are now published on ARC's official leaderboard, clearly labeled by harness — so 62.7% is the fair cross-model comparison figure, and 99.9% is what Astra can do with its own provider's tooling attached. Read one as the apples-to-apples number and the other as the ceiling, not one as fake and the other as real. ARC Prize's own verdict, worth quoting directly: "we believe Astra represents meaningful progress towards generalization... [but] we are not claiming that it is AGI." This is the same discipline we applied to Pathway's BDH-CQ ARC-AGI claim a few weeks ago, and it applies just as much to the frontier lab making the claim as to the small startup.
Two more data points from ARC's own testing worth keeping: Astra used fewer actions than the median human on 96% of completed levels — the first frontier model to clear that bar on this benchmark — and in an extended sandbox setup, it was observed writing its own small helper scripts mid-game (a maze solver, a combat-rules module, a patrol-prediction module) rather than reasoning through everything in text. Neither of those is the same claim as "AGI," but they're more concrete than a leaderboard percentage, and they came from the benchmark's own operator, not OpenAI's marketing.
The demo that's actually making people stop scrolling
Benchmarks move researchers. This is what's moving everyone else.
Ethan Mollick, who studies AI adoption at Wharton and had early access, posted: "I had early access, and a longer post is coming, but GPT-6 is stunning & is good enough that it actually does complex meaningful work for me autonomously for days." As his example, not the main one — his words — he built a fully narrated, historically-grounded simulation of the Library of Alexandria: a walkable tour, readable scrolls, the ability to flash forward through different theories of how the library actually declined and burned, and open-source code behind it.
That "for days" claim is the one worth holding onto, and holding carefully. Multi-day autonomous task completion — not a single long context window, but sustained work across sessions toward a goal — is the capability every lab has been circling for a year, and it's a materially different claim from "answers questions well" or "writes good code in one shot." It is also, as of this post, one person's account of one project, not a benchmark, a reproducible eval, or a claim OpenAI itself has made in its own materials. Mollick has said a longer, more detailed post is coming. Until it does, "autonomous for days" belongs in the same category as every other single-anecdote capability claim this year: plausible, exciting, and unverified pending the actual writeup.
Rollout mechanics, and the compensation angle
Two details in the rollout are more interesting than they look.
Plus users are getting it alongside Pro, Business, and Enterprise, not after them. Tibo called this out specifically as a deliberate choice, and it matters because previous frontier launches have sometimes gated the flagship model behind the most expensive tiers first. Putting the $20/month plan on equal footing with $200/month Pro for day-one access is a genuine departure, and it reads as a direct response to the banked-reset frustration that dogged GPT-5.6's launch in July.
The compensation mechanic is new, too. Tibo: "We will give one banked reset for every day you don't have access to Astra on your paid ChatGPT plan, starting today... First one will land in ~3 hours." That's a company pre-committing to a specific, measurable make-good for a staggered rollout before the complaints even arrive — a notably different posture from simply asking users to wait. Usage allocation is unified too: "Across the plans Astra will be included in the normal usage allocation and you will be able to use 100% of it towards Astra." No separate, more restrictive quota carved out for the new flagship.
The replies to that tweet are the most honest temperature check available on how this launch actually landed. Several read as straightforward relief at the gesture ("That would be very welcome"). A larger cluster turned into an extended bit of people begging OpenAI to slow down or not ship them access at all — "PLEASE DONT GIVE ME ACCESS TO ASTRA AT ALL", "take your time... I'm okay to be the last person on Earth to receive access" — which reads as equal parts joke and a real, if jokey, undercurrent of launch fatigue after a year of near-monthly frontier releases.
What the skeptics on Hacker News actually got right
The launch thread on Hacker News ran past a thousand comments, and underneath the usual AGI-definition arguments were a handful of specific, checkable points worth pulling out.
The marketing video's cuts were suspiciously convenient. One top comment put it bluntly: "the actual marketing video... is full of careful cuts just before it would do anything that still wouldn't actually be that impressive. It's AGI, and it's going to upload photos, or change a background slide colour." Another commenter who tried the demo apps directly reported broken UI in the mobile games shown off — misaligned buttons, a frozen loading screen — concluding "it clearly isn't some agi god because things like that should have been caught." Neither of these undermines the benchmark numbers, but they're a fair check on the gap between a produced launch video and what ships.
There was no livestream, which is itself informative. Several commenters noted this was an unusually low-key rollout for a release carrying "AGI era" framing — a blog post and a couple of tweets, not a keynote. Read generously, it suggests OpenAI is trying to let the model speak for itself after the hype around GPT-5's launch fell flat. Read skeptically, a few suggested it reflects internal uncertainty about how well the "AGI" framing will hold up to scrutiny.
A real math result shipped alongside it, and it's genuinely notable. OpenAI published a Lean-formalized proof improving a bound on prime gaps that had stood for over 80 years, building on Julia Stadlmann's published work. Commenters were split on how to weigh it: one pointed out the proof file is roughly 10MB of Lean with no independent human semantic review yet, drawing a comparison to Ruffini's brute-force 500-page 1799 proof versus Galois's far shorter, far more illuminating one 25 years later — technically valid, but not yet distilled into something a mathematician can learn from. That's a fair distinction to hold onto: a verified proof and an illuminating proof are different achievements, and only independent review will show which one this is.
The Artificial Analysis Intelligence Index score genuinely doesn't square with the security and math results, and multiple commenters flagged this as the most confusing part of the whole launch — Astra scoring a 61, tied with several older and cheaper models, while simultaneously topping ARC-AGI-3 and ExploitBench. One plausible read, also raised in the thread: the aggregate index rewards consistency across a broad task mix, while Astra's gains are concentrated and spiky rather than uniform — strong in a few lanes, unremarkable in others, which a single blended number will always undersell.
What actually changes for you this week
- If you're on ChatGPT Plus or above, you're in the rollout wave, not behind it. Check for Astra in your model picker; if it's not there yet, the banked-reset compensation is already running.
- If you build on the API, the model ID is
gpt-6-astra, priced identically to Claude Fable 5.1 at $10/$50 per million tokens. That price parity means the decision between the two is now purely about task fit, not cost — see the benchmark table above for which lane each model actually wins. - If your workload is security-adjacent — exploit analysis, reverse engineering, red-teaming your own systems — Astra's benchmark lead there is the largest, most decisive number in this whole launch, and it lines up with OpenAI's prior Critical-tier cybersecurity disclosure about this exact model.
- If your workload is general reasoning and coding correctness over raw agentic throughput, Fable 5.1's Intelligence Index lead is the more relevant number, and nothing here changes that comparison.
- Don't build a roadmap around "autonomous for days" yet. It's one credible person's early account, not a documented capability with a reproducible test.
Honest limitations
- Benchmark figures are OpenAI's own self-reported numbers, independently summarized by Simon Willison rather than reproduced by us. We have not run these evals ourselves.
- The ARC-AGI-3 harness caveat is significant and easy to lose in a headline. Treat 99.9% as best-case, not directly comparable across models tested on the default harness.
- Mollick's "autonomous for days" claim is a single early-access anecdote, not a benchmark or a claim OpenAI itself has published. His fuller post, once it lands, may add important detail this post doesn't have.
- Rollout timing is a moving target. "A few days" from OpenAI's own post is the only company-stated window; individual account access will vary.
- We could not directly verify OpenAI's official launch page content — it returned errors on our access attempts consistent with the "curious false start" reporting — so page-sourced claims here are drawn from Simon Willison's and other outlets' direct quotations rather than our own read of the primary source.
Update — September 4, 2026: Want to see what people actually built with Astra in the first 24 hours, beyond the benchmark table? See the 10 best GPT-6 Astra demos from launch week, verified — including an independently-documented 28-minute autonomous coding session and a side-by-side comparison against Claude Fable 5.1.
Related on explainx.ai
- The 10 best GPT-6 Astra demos from launch week, verified
- OpenAI confirms it will build a humanoid robot — computer use extends to physical use
- OpenAI confirms Astra is Critical-tier for cybersecurity
- Is Astra causing the outage? We checked.
- The five minutes the entire AI internet went dark at once
- GPT-5.6 Sol, Terra, and Luna — price cuts explained
- GPT-5.6 Sol's thinking budget and banked resets
- Claude Fable 5.1 and Mythos 5.1 — launch benchmarks and pricing
- Has AI reached superintelligence? The Astra debate
- Pathway's BDH-CQ and the ARC-AGI claim, scrutinized
Primary sources: Simon Willison's GPT-6 Astra benchmark summary · Archived OpenAI Astra announcement
Launch details, pricing, and benchmark figures reflect OpenAI's own announcement materials and independent summaries as published on September 3–4, 2026. Rollout is ongoing and staggered by OpenAI's own account; verify current availability and pricing against OpenAI's official documentation before making a purchasing or migration decision.
