New paper: Passing Lean checks doesn't prove AI math proofs are correct — faithfulness is uncomputable
rohanpaul_ai · x · 2026-10-08
A new arXiv paper, "Navier-Stokes lost in translation", argues that autoformalisation — translating natural-language proofs into Lean for machine verification, as used in OpenAI's announced Navier-Stokes blow-up proof — offers no guarantee the original argument is correct.
Key points:
- The authors demonstrate a chatbot turning an incorrect proof into a valid Lean proof by silently fixing the error during translation.
- Resolving ambiguities in mathematical NL text — required for semantically faithful translation — sits arbitrarily high in the Solvability Complexity Index (SCI) hierarchy (SCI = ∞), making it harder than any computational problem including the Halting problem (SCI = 1).
- Conclusion: no AI translator can always be semantically faithful, so a passing Lean check says nothing about the correctness of the original NL proof.
The paper is by Alexander Bastounis, Fabian Circelli, and Anders C. Hansen, and includes several practical examples of AI mistranslations of NL statements and proofs into Lean.
More from Research
- Are URM and Universal Transformers the forgotten architecture beating standard LLMs? — moschles · 2026-10-09
- Researcher: Use AI to Rewrite Machine-Generated Math Proofs Into Human-Readable Forms — jd_pressman · 2026-10-09
- AutoScientist's two-agent checklist loop auto-audits every training example — sarahookr · 2026-10-09
- NeurIPS GenAI4Health oral: retrieval-based medical fact-checking fails in ways bigger models can't fix — mdredze · 2026-10-09
- Bigger models, more reasoning, better sources won't fix medical fact-checking, researchers say — mdredze · 2026-10-09
- DEX best abstract: top LLMs catch many physician diagnostic errors, but big gaps remain — mdredze · 2026-10-09