2026-09-08

OpenAI's Own Writeup Walks Back the Navier-Stokes Overclaim — Its Account of the Buckmaster Dispute Still Ignores His Worst Allegations

AIScienceBusiness🌍 Global

OpenAI's tweet said it was "sharing a solution to the Navier-Stokes Millennium Prize Problem," produced by "a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra" — the first public acknowledgment that a model beyond Astra already exists internally. OpenAI's own fuller writeup, published alongside a public Lean 4 formalization on GitHub, is considerably more precise than the tweet compressed it into — and reading the two side by side is its own lesson in how much gets lost when a careful result becomes a headline.

What OpenAI's own writeup actually establishes

The writeup states the result resolves the Navier-Stokes Millennium Prize problem "by establishing statement 'C' (and also 'D') in the official Millennium Prize formulation" — that is, one of the four specific sub-statements the Clay Mathematics Institute itself defined as a valid resolution when it named the problem in 2000, not some easier side-problem outside the prize's actual scope. Concretely: an initially smooth, resting fluid, under a smooth external force, develops a singularity in finite time while its total energy stays finite throughout — established for both the whole space and the periodic case. That correction is worth making plainly, since it's more precise than treating "the equations had a forcing term" as automatically disqualifying; the official problem statement allows for exactly this kind of forced setup as one of its own four legitimate paths to resolution. Separately, and as a stepping stone along the way, OpenAI's agents also resolved the unforced Euler regularity problem — the inviscid limit of Navier-Stokes, with no external force at all — which OpenAI's own writeup says "surprised" the researchers running the effort.

What OpenAI does not claim is the money: "We do not intend to claim the Millennium Prize for this result," the writeup states directly, adding that "this is not a culmination, but rather a snapshot in time, of progress on AI development." That's a genuinely humble framing sitting right next to a headline-grabbing tweet that didn't carry any of it — and it's worth noting that the actual mathematical content still hasn't been independently reviewed by the field, let alone accepted by Clay, whose own rules require two years of published scrutiny before any prize is paid regardless of correctness.

The process behind it is described in detail, and it's worth taking at face value since it's checkable in outline: the effort began September 1 after OpenAI heard rumors that "two Millennium Prize problems had been resolved," which turned out to trace back to Alpöge and Buckmaster's own work. OpenAI's agents were split into groups and prompted with different variants of each open Millennium problem — for Navier-Stokes specifically, versions "A" and "B" (which would yield a proof of global regularity) and "C" and "D" (which would yield a disproof via singularity) were assigned to separate groups in parallel. A side-effort of nearly 100 agents spent about 50 hours resolving the unforced Euler case first; once that landed, OpenAI redirected resources toward Navier-Stokes specifically, cross-pollinating agent groups via Codex to consolidate promising partial results. The group that found the Navier-Stokes resolution ran on the order of 10,000 concurrent agents for about 88 hours, exchanging 2.7 million messages and roughly 130 billion output tokens — with a further 17 hours for Lean formalization and verification, run through GPT-6 Astra itself rather than the more capable internal model that found the proof.

The same week, in the open: a comparable result with none of the ambiguity

Independently, NYU Courant Institute mathematician Tristan Buckmaster — a Clay Research Award winner in his own right — and Anthropic researcher Levent Alpöge published three papers proving finite-time blowup with smooth forcing for the incompressible porous medium equation, the two-dimensional Boussinesq system, and the three-dimensional incompressible Euler equations, building directly on a multi-year research program by Diego Córdoba and Luis Martínez-Zoroa. The pair disclosed their AI usage in detail: Claude helped identify and reproduce key elements of the prior Córdoba–Martínez-Zoroa argument, and Claude alongside OpenAI's own Codex helped write the manuscripts. Every one of their results is publicly posted and formally verified in Lean. Terence Tao called the work "a remarkable achievement." OpenAI's own writeup confirms the two results are genuinely distinct — "our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)" — and explicitly recognizes "their priority" on forced Euler.

Two accounts of the same contact, and one of them leaves out the damaging part

OpenAI's writeup describes reaching out to Buckmaster and Alpöge only after completing its own Navier-Stokes work and Lean verification on September 6, "to offer a concurrent release of our result and to recognize their priority in a joint announcement," offering them visibility into its prompts and eventually the proof itself. It's a tidy, collegial account — and it is also, conspicuously, not the account Buckmaster gave the same day. Buckmaster's own public statement says Sébastien Bubeck, who leads OpenAI's mathematical research, twice pushed to have Alpöge dropped from authorship specifically because he works at Anthropic, and describes two concrete proposals: either Buckmaster and Alpöge post their Euler result while OpenAI posts its Navier-Stokes result the next day, or Buckmaster alone write up the Navier-Stokes result crediting an internal OpenAI model — without Alpöge at all. Buckmaster says he declined both, said he'd go public if OpenAI proceeded as proposed, and recounts being asked, "Why would you ruin your career?" and then told, "If you don't want me to be nice, then I don't have to be nice." Bubeck has separately called Buckmaster's characterization "false and inflammatory" and said he acted "following academic norms," without addressing those specific claims point by point. OpenAI's official writeup doesn't address them either — it describes a courteous outreach to "recognize priority," and simply never engages with the pressure, the two proposals, or the quotes Buckmaster made public. Silence on the specific allegations, in a document otherwise detailed enough to name exact hour counts and token totals, is itself worth noticing.

OpenAI's own tweet still reads like a pre-emptive denial

Read against that backdrop, the second post in OpenAI's own tweet thread is more revealing than it looks in isolation: "We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models. However, our proofs differ significantly." That's a specific denial of a specific suspicion — improper visibility into Buckmaster and Alpöge's unpublished work — volunteered unprompted, in the same breath as congratulating them. Combined with Buckmaster's account of being pressured over authorship in the days before either side published, the denial reads less like routine caution and more like OpenAI addressing a question it expected to be asked.

This isn't OpenAI's first math-credibility problem this year

This blog covered the last one in detail: last October, an OpenAI VP tweeted that GPT-5 had "found solutions to 10 (!) previously unsolved Erdős problems," which turned out to be false — the model had mostly located existing literature — and the tweet was deleted after mathematician Thomas Bloom called it "a dramatic misrepresentation." This week's tweet ("a solution to the Navier-Stokes Millennium Prize Problem") makes the same compression error in miniature: technically defensible once you read the full writeup's careful "statement C and D" framing and its explicit disclaimer about not claiming the prize, but stripped of all of that nuance in the version most people actually saw. OpenAI's response to last October's embarrassment, this blog noted at the time, was visibly building a habit of getting the field's specific skeptics on record before a headline claim landed — Fields Medalist Timothy Gowers on a May result, Bloom himself on a batch of ten results in July. This episode inverts that pattern: instead of securing endorsement before going public, OpenAI's own math lead is accused of pressuring a mathematician into a credit arrangement that would have erased his coauthor, and the company's own published account of that contact simply omits the allegation rather than rebutting it. The Leiden Declaration on AI and Mathematics — signed by more than a thousand mathematicians including Tao, and cited by OpenAI itself in July — warns specifically about unreliable results, missing attribution, and exaggerated claims. This week's episode brushes against all three: a result not yet independently reviewed, a live authorship dispute OpenAI's own account doesn't engage with, and a tweet meaningfully less careful than the writeup backing it.

What to expect next

  • Watch for Bubeck's promised fuller response, and specifically whether it addresses the two proposals and the "ruin your career" exchange directly, or continues past them the way OpenAI's own writeup already has.
  • Watch for independent mathematical review of the actual proof, not just the Lean certificate's internal consistency. OpenAI's own writeup is unusually precise about which official Clay sub-statement it targets — the open question is whether outside mathematicians confirm the argument holds, not whether it's aimed at the right target.
  • Watch the Clay Mathematics Institute's own position, given OpenAI itself isn't seeking the prize. Two years of published acceptance is required before any award regardless, so nothing about the $1 million resolves quickly either way.
  • Watch whether Anthropic comments on how its own researcher was treated. Alpöge is a named party in a credit dispute with a direct rival lab, and neither OpenAI's tweet nor its official writeup addresses the specific pressure he and Buckmaster describe; Anthropic's silence or response is itself a data point.