2026-07-31

OpenAI's Math Claims Have a Credibility Problem. This Time They Might Have Fixed It.

AIScience🌍 Global

Nine months ago, an OpenAI VP tweeted that GPT-5 had "found solutions to 10 (!) previously unsolved Erdős problems." It hadn't — the problems were listed as open only in the sense that mathematician Thomas Bloom, who maintains the tracking site erdosproblems.com, was personally unaware of a published solution; the model had mostly located existing literature. Bloom called the tweet "a dramatic misrepresentation," OpenAI's own researcher admitted the error, the tweet was deleted, and Demis Hassabis called the whole episode "embarrassing." That history matters for reading what OpenAI published this week: ten new results on problems that have seen no progress in a decade or more, produced by an internal, unreleased version of a model called Astra. The claim is nearly identical in shape to the one that blew up in October. The reason to take this one more seriously is that Thomas Bloom is the one vouching for it.

What's actually being claimed

Ten results, spanning high-dimensional geometry, coding theory, arithmetic circuit complexity, group theory, operator algebras, quantum complexity, lattice cryptography, and extremal combinatorics — among them new upper bounds on sphere-packing density down to the Cohn–Elkies threshold, a disproof of Connes's rigidity conjecture (that certain groups are uniquely determined by their von Neumann algebras), a new arithmetic-formula lower bound for computing the permanent, polynomial-factor hardness of approximation for the closest vector problem (a foundational question in post-quantum cryptography), and a resolution of Erdős problem 183 — a superexponential lower bound for multicolor triangle Ramsey numbers. OpenAI's own framing is specific about the bar: problems with no progress on the main result "for at least a decade, and in most cases much longer."

The production process is worth noting on its own: an internal version of Astra generated the mathematical arguments — OpenAI estimates the total token cost at roughly $2,000 at Sol API rates for all ten — humans then worked with the same model to prepare the arguments into manuscripts, and the model formalized each proof into a Lean certificate, a machine-checkable formal verification. OpenAI is also releasing the model's own narration of its reasoning process for each result.

Why the Lean step is the part that matters most

A natural-language proof can be persuasive and wrong at once — mathematicians have been burned by AI-generated arguments that read fluently but contain a gap a careful reader would eventually find. A Lean certificate is a different kind of claim: it's checked by a proof assistant, not a reader's attention span. That's the same instinct behind Google's ScientistOne / Chain-of-Evidence framework, published one day before this OpenAI post — the trust problem in AI-generated research output isn't "is the model smart enough," it's "can the claim be checked independent of whoever is making it." OpenAI's own attribution language leans the same way: "claiming human authorship for a proof generated entirely by an AI system would misrepresent both the system's contribution and the nature of genuine human intellectual work. We helped prepare the manuscripts and formalize the proofs in Lean, and we take responsibility for their correctness, while the mathematical arguments themselves were generated by our system." That's a more precise division of labor than most AI-and-science announcements bother to draw.

The May precedent, and the credibility test this time

This isn't OpenAI's first attempt at this exact claim shape. In May, the company shared an AI-generated disproof of the Erdős unit-distance conjecture, found while evaluating an unreleased model — and this time, the response from mathematicians who'd have every reason to be skeptical was strongly positive. Fields Medalist Timothy Gowers said that if a human researcher had submitted that paper to Annals of Mathematics, he'd have recommended publication "without any hesitation." On this week's ten results, Bloom — the same person who caught the October overclaim — called them "big news," and specifically ranked them above the May disproof in significance: "Maybe not bigger than a proof of unit distance would have been, but in terms of constructions, this is big."

That's the actual signal worth reading here, more than the results themselves: the skeptic with the most reason to distrust OpenAI's math claims, and the specific track record of having debunked one, is on record finding this one credible. Track records compound in both directions.

The caveat that shouldn't get lost in the headline

Independent analysis of the May result found something worth keeping in view for this batch too: the model didn't invent new mathematical techniques from nothing — it recombined and extended existing ideas drawn from multiple subfields, which is a real and difficult skill, but a different one than originating a new proof method. Human mathematicians were also involved in tightening and extending the May result after the model's first pass, and the same collaborative structure is explicit in this week's release — humans prepared the manuscripts alongside the model, they didn't just publish raw model output. None of that makes the results less real. It does mean the honest framing is "a system that can search an enormous space of existing mathematical technique faster and more broadly than a human can," not "a system that reasons about mathematics the way a human does." OpenAI's own post explicitly credits the signers of the Leiden Declaration on AI and Mathematics — the June 2026 statement, backed by more than a thousand mathematicians including Terence Tao and Peter Scholze and endorsed by the International Mathematical Union, warning specifically about unreliable results, missing attribution, and exaggerated claims — for raising exactly these questions. Naming that document by name, in a post making exactly the kind of claim it warns about, reads like an attempt to hold the distinction rather than blur it for a better headline.

Where this sits against the rest of the science vertical

This is the same zero-dollar-capital, all-cost-curve vertical this blog has tracked in formal mathematics via Mistral's Leanstral — except it's arriving from a frontier general-purpose lab rather than a dedicated math model, using an unreleased flagship rather than a specialized release, and landing on problems open for decades rather than benchmark problem sets. It also arrives two days after OpenAI's own ChatGPT for Academic Researchers program went live for 100,000 scientists and mathematicians — this post reads, in part, as that program's proof of concept, published before most of the 100,000 have even gotten their accounts.

What to expect next

  • Formal verification becomes the minimum bar for any AI mathematics claim, not a nice-to-have. After October, "the model said so" isn't enough from OpenAI specifically; a Lean certificate is the only version of this announcement that survives scrutiny on arrival.
  • Named mathematician endorsement becomes part of the release strategy. Gowers on the May result, Bloom on this one — OpenAI is visibly building a track record of getting the field's specific skeptics on record before the headline lands, which is a lesson learned in public.
  • Watch which subfields get the next wave. Ten results across eight areas in one announcement suggests Astra's math capability isn't narrow — the interesting question is whether the next ten look like more of the same recombination-at-scale or start to include genuinely novel technique.

References: OpenAI — "Ten advances in mathematics and theoretical computer science" · the-decoder.com — Astra's ten math solutions · Scientific American — the May Erdős unit-distance disproof · TechCrunch — the October 2025 "embarrassing math" retraction · Leiden Declaration on Artificial Intelligence and Mathematics · related coverage on this blog: ScientistOne / Chain-of-Evidence · The Science Vertical