FULL STORY

OpenAI Math Proofs Under Fire: Retractions and Lean Verification

A Cambridge paper challenged the reliability of Lean verification of OpenAI's proofs, prompting OpenAI to retract three manuscripts from its new math repo. Researchers then built on the work with Lean-verified improvements.

2026-10-08 ~ 2026-10-09 · 3 episodes · 21 posts

Episode 1 · Cambridge paper argues Lean verification does not certify proofs, targeting OpenAI's Navier-Stokes claim (2026-10-08, 14 posts)

Cambridge researchers Bastounis, Circelli, and Hansen published a paper on arXiv titled "Navier-Stokes lost in translation" (numbered 2610.08144), systematically challenging the practice of endorsing mathematical proofs via "AI autoformalization + Lean mechanical verification," and directly questioning OpenAI's previously announced blow-up proof for the Navier-Stokes equations.

Confirmed

  • The paper's core claim: automatically formalizing a natural-language proof into Lean and passing Lean's checks does not show that the original natural-language proof is correct.
  • The authors present a concrete case: a chat model "quietly fixed" an incorrect proof during translation into a valid Lean proof, which Lean still accepted—showing the translation process can mask errors in the original proof.
  • The paper proves theoretically that the faithful-translation problem is undecidable, and that translating ambiguous natural language is even harder than the halting problem.
  • The paper explicitly targets OpenAI's claimed Navier-Stokes proof, arguing its formal verification pipeline has fundamental flaws.
  • Pedro Domingos, Rohit Paul, and other bloggers shared the paper, emphasizing its impact.

Unconfirmed

  • Regarding whether "GPT-6 fixed Peter Bel's erroneous Navier-Stokes proof," Elliot Glazer remains cautious: he leans toward the view that the Lean formalized proof and the English paper do not correspond one-to-one, further undermining the credibility of "deformalization"—but this is his personal judgment, not a paper conclusion.

Why it matters

  • If passing Lean verification cannot retroactively guarantee the original proof's correctness, the credibility of AI mathematical proof claims needs to be re-evaluated, and the role of formal verification in AI math workflows faces a fundamental challenge.

Episode 2 · OpenAI Retracts Three Hodge-Conjecture Papers Over Sign Error, 42% of Results Now Formalized (2026-10-08, 5 posts)

OpenAI published its first changelog just two days after making the openai/math repository public: 3 manuscripts retracted, 14 revised, and 13 citations updated, bringing the total number of manuscripts down from 722 to 719. The retraction stemmed from a notation error that invalidated a stability-trace cancellation argument, and two other papers relying on that argument were retracted as well; according to 机器之心, all three papers relate to the Hodge conjecture, touching on topics such as the algebraicity of Weil classes and K3 surfaces.

Confirmed

  • Around October 7, OpenAI retracted 3 papers and revised the other 14 manuscripts.
  • The repository added 6 Lean formalizations and 19 modifications.
  • Of the 719 core (headline) results, roughly 300 — about 42% — have been formalized in Lean, and the team says updates will continue.

Why it matters

  • The episode shows formal verification is becoming a credibility gatekeeper for AI-generated mathematics: the error was exposed by the Lean formalization process, prompting a timely retraction.
  • A 42% formalization coverage means most results have yet to be machine-verified, so the repository's credibility still depends on ongoing formalization work.

Episode 3 · Researchers Improve OpenAI's Math Proof, Verified in Lean (2026-10-08, 2 posts)

Researchers claim significant improvements on OpenAI's recent mathematical proof results, verified in the Lean formal proof system, with code open-sourced on GitHub.