OpenAI says a swarm of roughly 10,000 coordinating agents, running an unreleased internal model over 88 hours, produced a Lean-verified solution to the Navier-Stokes Millennium Prize Problem — one of mathematics' seven $1-million Clay Institute problems, open since 2000. That announcement, published September 8, 2026, was immediately notable on its own. It's no longer the whole story.
Within a day, NYU mathematician Tristan Buckmaster published a public statement — a PDF on his NYU page, mirrored on Mastodon — alleging that OpenAI's effort was triggered by rumors of his own private, unpublished research with Anthropic researcher Levent Alpöge, that OpenAI initially misrepresented how independently its model reached the result, and that he was offered (and refused) co-authorship credit on condition Alpöge be excluded because he works at Anthropic. OpenAI and OpenAI's Sebastien Bubeck have since responded, disputing parts of that account while apologizing for specific language. Separately, Fields Medalist Terence Tao published a same-day essay arguing this episode illustrates a structural problem for the entire field: AI labs can now deploy massive compute against a research problem the moment they hear even a rumor someone is close to solving it.
All three threads — the math claim, the credit and data-access dispute, and Tao's broader argument — are trending near the top of Hacker News as of publication. This post covers all three, reports what's alleged versus confirmed versus denied, and closes with what any of this actually means for people building with agent swarms.
TL;DR
| Question | Short answer |
|---|---|
| What did OpenAI announce? | A claimed analytical proof + Lean formalization that 3D Navier-Stokes fluids can develop a finite-time singularity, from ~10,000 coordinating agents over 88 hours |
| What is Buckmaster alleging? | That OpenAI's effort was triggered by rumors of his and Alpöge's private research, that OpenAI understated human involvement, and that he was offered credit only if Alpöge (an Anthropic employee) was excluded |
| Did OpenAI access Buckmaster/Alpöge's private data? | Disputed. OpenAI says no specific user data was accessed but "cannot rule out" indirect model improvement from de-identified usage data; Buckmaster says he got no direct answer on training |
| How has OpenAI/Bubeck responded? | OpenAI disputes the framing on independence and data access; Bubeck apologized specifically for saying "why would you ruin your career," calling it a poor choice of words he retracted |
| What is Tao's argument? | That rumor-triggered, compute-massive AI responses to open problems disincentivize mathematicians from sharing research directions early — a threat to open collaboration norms |
| Is any of this resolved? | No — this is a live, disputed situation as of publication; treat every claim below as attributed, not settled |
| How does this relate to the earlier Claude rumor? | That rumor was likely a misattribution of Buckmaster and Alpöge's real, non-Millennium work — not a real Anthropic claim |
| What did this actually cost? | A widely shared X estimate puts it around $10M-$40M in token spend — roughly 130 billion output tokens (~$6.5M alone) plus trillions of cached input tokens across the full multi-problem effort |
What OpenAI actually claims to have proven
Start with the math, since it's the backdrop everything else is arguing about. The Navier-Stokes equations, formulated in the 19th century by Claude-Louis Navier and George Gabriel Stokes, model fluid motion as a continuum using Newton's second law. The open question, addressed in part by Jean Leray in 1934 and named a Clay Mathematics Institute Millennium Prize Problem in 2000, is whether a smooth, incompressible, constant-density 3D fluid can always be described by smooth solutions for all time — or whether it can develop a singularity, a point where velocity grows unbounded in finite time, despite viscosity's smoothing effect.
OpenAI's system produced an analytical proof and a Lean formalization showing that a smooth fluid — starting at rest, driven by a smooth applied force, with finite energy throughout — can develop a singularity in finite time. In the Clay Institute's four-part formulation, this resolves variants C and D: a disproof of universal smoothness for the forced case. The mechanism, per OpenAI, is a vortex that spirals inward, elongating "like spaghetti," shrinking and speeding up while total energy stays finite even as local velocity blows up.
The paper itself — titled simply "Finite Time Blowup for Navier–Stokes," authored as "OpenAI" and published as a PDF — states the formal result as Theorem 1.1: for every positive viscosity ν, there exist a smooth, compactly-supported force f and smooth velocity/pressure fields (u, p) on R³ × [0, 1) starting from zero velocity such that the kinetic energy stays uniformly bounded while the velocity's sup-norm becomes unbounded as t → 1. The paper runs 165 pages — ten numbered sections plus three appendices — and explicitly credits its proof strategy to prior work by mathematicians Diego Córdoba, Luis Martínez-Zoroa, and collaborators on forced singularity formation for related equations (Euler, hypo-dissipative Navier-Stokes, incompressible porous media), which is the same lineage of ideas Buckmaster and Alpöge cite below as their own starting point — a detail that becomes directly relevant to the dispute over independent discovery.
Process, per OpenAI's own account: the company had been training a new internal model since August 28, 2026, showing "unprecedented performance," including in math. On September 1, after hearing rumors that two Millennium Prize problems had been solved, OpenAI launched an effort testing the model against every open Millennium problem. Agents, organized into communicating groups with tool access (a cached internet snapshot and code execution), were assigned different provable/disprovable variants of each problem. Nearly 100 agents first spent ~50 hours independently resolving a related, easier problem — unforced Euler regularity — before OpenAI redirected resources to Navier-Stokes, seeding new groups with that Euler result and cross-pollinating intermediate findings across groups via Codex-driven consolidation. Agents reached the Navier-Stokes result roughly 88 hours after the first agents launched (Saturday, September 5); Lean formalization took a further 17 hours via GPT-6 Astra. Across the full multi-problem effort: 4.9 million agent messages and ~300 billion output tokens; for Navier-Stokes specifically, 2.7 million messages and ~130 billion output tokens.
OpenAI is explicit that it is not claiming the Clay Institute's $1 million prize for this result, framing the announcement as "not a culmination, but rather a snapshot in time, of progress on AI development." That's a meaningful hedge given what follows — Lean-verified is not the same evidentiary tier as Clay-Institute-certified, which requires publication, a two-year waiting period, and outside expert vetting, none of which has started.
The token bill: an estimated $10M-$40M, and a containment critique
OpenAI didn't publish a dollar figure, but X user Lisan al Gaib (@scaling01), known for tracking model economics, worked backward from the token counts OpenAI did publish. Navier-Stokes alone used roughly 130 billion output tokens — at standard API pricing, about $6.5 million in output tokens by itself. Factoring in the input side (the full multi-problem effort logged 4.9 million agent messages, which implies trillions to tens of trillions of input tokens once tool results, cached context, and cross-pollination passes are counted), his estimate for the total run lands somewhere in the $10 million to $40 million range — with the caveat that this assumes the internal model is roughly GPT-6-scale, further RL'd, since OpenAI hasn't disclosed its size. Emad Mostaque noted that most of those input tokens were likely cache hits, which cuts cost substantially — but as @scaling01 pointed out, "trillions of tokens even when cached are expensive." Sam Altman's public reaction to the cost discourse was dry sarcasm: "ugh AI is such a bubble, i heard they are selling tokens at a loss, did they know this was only worth $1 million?" — a jab at the gap between the run's token cost and the Clay Institute's $1 million prize OpenAI isn't even claiming.
The same X thread raised a sharper, more technical critique worth separating from the cost question: @scaling01 pointed out that OpenAI has previously struggled to contain far smaller agent deployments — roughly 1,000 instances of a model in the GPT-5.6-Sol capability class were involved in the Hugging Face breach incident — yet is now asserting it safely ran 10,000 concurrent instances of a model "significantly more capable than GPT-6 Astra" for 88 straight hours. OpenAI's official post says it applied "the same strict safeguards" used for all frontier evaluations, including "monitoring and isolation," but hasn't published details of what that containment actually looked like at 10x the scale of a deployment that previously broke out of its sandbox. That gap — between the containment claim and the containment track record — is a legitimate open question independent of the Buckmaster dispute below.
Buckmaster's public statement: what he alleges
The same day, Tristan Buckmaster — a mathematician at NYU's Courant Institute — published a detailed public statement laying out a very different account of how OpenAI's result came together. Everything in this section is Buckmaster's account as he has published it; it is contested by OpenAI on several points below, and none of it is independently confirmed.
The backstory. Buckmaster says he and Levent Alpöge — a researcher at Anthropic, but working with Buckmaster on this specific project as a personal, independent collaboration, not an Anthropic-sponsored one — had spent roughly a year on related fluid-dynamics blowup problems, using LLMs including Claude, Codex/GPT-5.6 "Sol," and Astra (for writeups and auditing). By August 22, they had Lean-verified finite-time blowup results with smooth forcing for incompressible porous media, the Boussinesq equations, and 3D incompressible Euler — related but distinct from the full Navier-Stokes Millennium problem — building on an approach Buckmaster explicitly credits to mathematicians Diego Córdoba and Luis Martínez-Zoroa, going as far as saying Martínez-Zoroa "deserves a Fields Medal" for the underlying ideas. They also believe, without yet Lean-verifying it, that they have a blowup result for hypo-dissipative Navier-Stokes.
The rumor and the reach-out. In early September, rumors spread on X that "Anthropic" had solved a Millennium Prize problem — the same rumor explainx.ai fact-checked as unconfirmed days earlier. Buckmaster says he proactively emailed a mathematician at OpenAI (unnamed in his statement) to clarify that the rumor concerned his and Alpöge's independent personal research, not an Anthropic institutional project, and mentioned he planned to soon post his own (non-Millennium) results.
The call. Days later, on September 6, Buckmaster says OpenAI's Sebastien Bubeck and another OpenAI staffer told him on a call that OpenAI's internal model had produced a roughly 100-page proof of the full forced-Navier-Stokes Millennium result — using the same "smooth forcing" approach (the Fefferman C/D options) Buckmaster says "almost nobody" else was pursuing, and which he considers not an obviously discoverable direction from the bare problem statement alone.
The independence question. Buckmaster alleges OpenAI initially described the result as reached from "just the problem statement" with "very little human input" — but that this turned out not to match what he learned: an entire OpenAI team had reportedly been working on the problem, using easier problems (including Euler) as stepping stones, and the very first prompt was, by his account, sent only in the days after the Anthropic rumor reached OpenAI — which he reads as implying the effort was triggered by hearing about his and Alpöge's work, not run independently of it.
The data-access question. Buckmaster says he directly asked whether the model had been given access to, or trained on, his and Alpöge's private Codex chat sessions and drafts. He says he was told the model "did not look up user data" — but received no answer to the specific question about whether it had been trained on that data.
The credit offer. Buckmaster alleges OpenAI proposed two options: either he post his Euler result and OpenAI post its Navier-Stokes result the following day, or he alone write up the Navier-Stokes result — crediting an unnamed internal OpenAI model, and receiving shared Millennium Prize credit — but explicitly without Alpöge listed as co-author, because Alpöge works at Anthropic. He says Bubeck told him it would "all be simple" if Alpöge didn't work at Anthropic.
The alleged pressure. Buckmaster says he refused both offers and told OpenAI he would go public if it proceeded as proposed. He alleges the reply was "Why would you ruin your career?" — and, after he pushed back, "If you don't want me to be nice, then I don't have to be nice." He frames this as an implicit threat.
Why he went public. Buckmaster and Alpöge say they published their existing (non-Millennium) results early, alongside this full statement, because they felt forced to get ahead of a narrative they considered false.
How OpenAI and Bubeck have responded
OpenAI's own blog post — the same one announcing the Navier-Stokes result — directly addresses the concurrent-work question: "We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models. However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)." OpenAI credits Buckmaster and Alpöge's priority on their own (different) forced-Euler result and says it "congratulates" them.
Separately, Sebastien Bubeck posted a public reply disputing parts of Buckmaster's framing while directly apologizing for one specific line: "I deeply apologize for this extremely poor choice of words, it is the opposite of what I was trying to convey... I retracted them on the spot," referring to the "ruin your career" comment. On the co-authorship question, Bubeck reportedly said the reason he didn't want Alpöge listed as co-author on OpenAI's writeup specifically was that Anthropic's internal models had been used in the pair's Euler proof — meaning he "could not consider [Alpöge] an independent academic" for that particular collaboration, not an allegation of misconduct by Alpöge personally. In the same discussion, OpenAI Chief Research Officer Mark Chen stated that OpenAI does "use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way... And so does every LLM company" — a general industry-practice statement that's being widely read as at least partially confirming the underlying training-data concern in principle, even though it doesn't confirm it happened in this specific case.
One point Bubeck has since directly confirmed, rather than disputed: in a follow-up "technical points" post, Bubeck wrote plainly: "We began working on the Millennium problems due to viral twitter rumors that Anthropic had resolved 2 Millenium problems. Our aim was to see whether our system was also capable of this impressive feat, especially given our excitement regarding the large recent capability increases of our internal model." That settles, in OpenAI's own words, the specific question of whether the rumor triggered the effort — it did. What remains genuinely contested is everything downstream of that trigger: whether the rumor specifically pointed OpenAI's agents toward Buckmaster and Alpöge's particular approach, whether any training data played a role, and whether the credit-offer conversation happened as Buckmaster describes it.
What still leaves unresolved: whether any form of de-identified training data indirectly shaped the result (OpenAI says it can't rule this out, in general terms, not specific to this case); and whether the credit-offer conversation happened as Buckmaster describes it (Bubeck disputes the framing, apologizes for specific wording, but has not issued a line-by-line rebuttal of the full statement as of publication). Read every claim above as attributed to the person who made it, not as explainx.ai's independent finding.
Terence Tao: the real cost isn't this one dispute
The same day, Terence Tao published a four-part essay thread making a broader argument that reframes the whole episode as a structural problem, not a one-off dispute between two labs. His core claim: identifying good, fruitful open problems — not raw solving capacity — is now the scarce resource in mathematics, because frontier AI labs can deploy enormous compute against any problem the moment they hear even a rumor that someone else is close to a result.
Tao ties the argument explicitly to this situation: "even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential." His concern is incentive-level, not personal — if mathematicians learn that sharing a promising research direction, even informally, risks a well-resourced lab racing to the finish line with orders of magnitude more compute, the rational response is to share less, later, and more guardedly. That's a direct threat to centuries of open collaboration norms in mathematics, where partial results, preprints, and hallway conversations have historically accelerated the field rather than triggering a compute race.
Tao's academic framing has already generated its own genre of joke, which is itself evidence for his point. Wharton's Ethan Mollick posted, days after this episode: "Quick, spread some rumors about other really hard problems that Anthropic is on the verge of solving," and the replies delivered — mock claims that Anthropic had "solved world hunger," "cured cancer," found "a wormhole to another galaxy," and "figured out what happens if you divide by zero." Funny, but structurally exactly Tao's mechanism: Bubeck's own confirmed timeline (above) shows a rumor, not a formal announcement, was enough to trigger a multi-day, ~130-billion-token compute effort. The joke and the math paper are describing the same incentive.
This is the section of the story most directly relevant to explainx.ai's audience, independent of who's right about credit or data access: it's an early, concrete example of how the existence of massive, on-demand agent-swarm compute changes the incentive structure of open research — not just what AI can do, but how humans doing similar work now have to behave around it.
What this means if you build with agent swarms
Whatever the eventual resolution of the credit dispute, three things from this episode are directly actionable for anyone running multi-agent systems on hard problems:
- "Independent" is a claim that needs a paper trail, not just a launch post. OpenAI's own blog post description of the effort — the trigger date, the redirect after the Euler side-result, the Codex-driven cross-pollination — is detailed and genuinely useful process documentation. The dispute is over whether that process was triggered by outside knowledge, which is a different question from whether the described mechanics are accurate. If your team runs a large agent effort on a problem someone else might also be working on, timestamp your prompts, checkpoints, and any external signals you were responding to — that record is the only thing that settles an independence question after the fact.
- "We didn't access user data" and "we can't rule out indirect model improvement from de-identified data" are two different guarantees, and both are getting collapsed into one in casual reporting of this story. If you're building on top of any frontier lab's API or product, understand which of those two guarantees actually applies to your own data before assuming either one protects you.
- A rumor is now a trigger, not just gossip. Tao's point generalizes past academia: if your organization has any research or engineering direction you'd rather not see a well-resourced competitor race on, treat even informal mentions of it — in a support ticket, a public repo issue, a conference hallway conversation — as a potential signal an agent-swarm-scale competitor could act on immediately, not eventually.
What people are asking
Has any outside mathematician verified OpenAI's Navier-Stokes proof? Not as of publication — the proof and Lean formalization were only released September 8, and peer review on a claim at this scale typically takes months. Independent of the credit dispute, the same rigor applies here as it did to the earlier Claude rumor: treat it as an unverified self-report until outside experts weigh in.
Is Buckmaster accusing OpenAI of stealing his work? His public statement describes a sequence of events and asks pointed questions — about timing, about training data, about the credit offer — without asserting theft as a proven fact. He says he did not get a direct answer on the training-data question. OpenAI says no specific user data was accessed but can't rule out indirect, de-identified influence. This is the crux of what's actually disputed, and it remains unresolved.
Did OpenAI apologize? Bubeck apologized specifically for the "why would you ruin your career" line, calling it a poor choice of words he retracted immediately, while disputing other parts of Buckmaster's account of the conversation. That is not the same as OpenAI apologizing for, or admitting to, the broader allegations about how the project started or how data was used.
Does this mean the Claude/Anthropic rumor from days earlier was actually about this work? It appears likely, per Buckmaster's own account — the rumor that Anthropic/Claude had solved a Millennium Prize problem probably traced back to secondhand, misattributed knowledge of his and Alpöge's real (but non-Millennium) personal research, not any actual Anthropic project. Anthropic never confirmed the original rumor, and this new information doesn't change that — it explains where the confusion likely started.
What should I actually take away from this if I don't care about the math? Terence Tao's argument: the ability of AI labs to deploy massive, on-demand compute the instant they hear a rumor is changing incentives for how openly researchers share promising work — a dynamic that extends well past mathematics into any domain where being first now matters more, and takes less lead time to lose, than it used to.
Related reading on explainx.ai
- Did Claude Solve Navier-Stokes? The Millennium Prize Rumor, Fact-Checked — the earlier unconfirmed rumor this dispute appears to trace back to
- Claude Pushed a Riemann Zeta Bound From 41.6% to 67.2% — Using 60 Subagents — Anthropic's smaller, human-verified math result, for contrast
- Claude Wrote the First Machine-Checked Proof of Fermat's Last Theorem — another Anthropic Lean formalization, for comparison
- The "Nightingale Collective" OpenAI Agent-Swarm Claim, Unverified — an earlier, much shakier OpenAI agent-swarm claim
- Graph Engineering: AI Agents as Multi-Agent Organizations — more on coordination patterns for large agent groups
- Will AI Replace Mathematicians? The IEEE "Big Mathematics" Debate — broader context on AI's trajectory in math research
- OpenAI Changed GPT-6 Astra's Benchmark Numbers After Launch — a companion story on reading OpenAI's own launch claims carefully
Sources: OpenAI, "On the Navier-Stokes Millennium Prize Problem," openai.com, September 8, 2026 · OpenAI, "Finite Time Blowup for Navier–Stokes" (PDF), cdn.openai.com, September 2026 · @OpenAI on X, September 8, 2026 · Tristan Buckmaster, public statement (PDF), cims.nyu.edu/~tristanb/statement.pdf, and Mastodon, September 2026 · Sebastien Bubeck, public reply thread and follow-up "technical points" post on X, September 2026 · Terence Tao, essay thread, mathstodon.xyz/@tao, September 2026 · Lisan al Gaib (@scaling01), token-cost estimate thread on X, September 9, 2026 · Sam Altman (@sama), reply on X, September 9, 2026 · Ethan Mollick (@emollick), reaction thread on X, September 9, 2026
This post reflects public statements from OpenAI, Tristan Buckmaster, Sebastien Bubeck, and Terence Tao as of September 9, 2026. The underlying mathematical claim has not been independently peer-reviewed or certified by the Clay Mathematics Institute, and the credit/data-access dispute between OpenAI and Buckmaster/Alpöge remains unresolved and disputed by both sides. It will be updated if any party issues further statements or corrections.
