Mathematician Daniel Litt Audits 19 of His Own Papers with AI, Finds 97.7% of Comments Flag Real Issues
Mathematician Daniel Litt publicly shared a full experiment on August 27: using the AI tool Refine.ink together with ChatGPT to automatically review 19 of his own papers and generate errata. Of 266 detailed comments, 239 were entirely correct, 21 partially correct, and only 6 wrong—meaning 97.7% pointed to real issues, including about 5 serious problems that could affect main results.
Confirmed
- The review covered 19 papers and produced 266 comments: 239 correct, 21 partially correct, 6 wrong, for a 97.7% rate of identifying real problems.
- The 5 serious issues were mostly missing technical hypotheses (e.g., projectivity or reducedness assumptions missing from main theorems). Litt believes careful readers could have inferred these implicit assumptions, but they do need to be fixed.
- Most findings were minor: typos, ambiguous phrasing, and the like.
- Litt also used AI to generate semi-automated errata, which he candidly calls "slop"—only lightly edited by hand. Spot checks showed they were largely correct but often too aggressive (e.g., replacing whole paragraphs instead of adding a single word). He has published them as-is to show what a few hours of effort can achieve.
- On the root cause of the errors, Litt reflected that in his first long paper nearly all mistakes stemmed from failing to propagate edits correctly throughout the text while polishing results—a classic way papers go wrong.
- Refine is not complete: it missed an error in a lemma of one of Litt's published papers, previously discovered by Jordan Ellenberg and Alex Smith, who published an erratum; it did not affect the main results.
Why it matters
- This is a systematic AI audit by a mathematician of his entire body of papers, offering quantified data (97.7% valid findings) and showing that AI review is already practically useful at catching genuine technical flaws.
- The experiment also exposed limitations: AI missed known errors and produced overly aggressive errata, so human verification remains indispensable.
2026-08-27 ~ 2026-08-27 · 8 related posts
Primary sources
- [source] AI audit flags 5 serious math issues, incl. missing theorem hypotheses — littmath · 2026-08-27
- Mathematician Litt: Paper Errors Mostly From Badly Propagated Edits, Not Deep Flaws — littmath · 2026-08-27
- Mathematician's AI-generated errata are "slop" but mostly correct — littmath · 2026-08-27
- A few hours of effort yields AI-generated errata for a paper archive — littmath · 2026-08-27
- [source] AI audit of 19 papers: 97.7% of comments identified real issues — littmath · 2026-08-27
- Mathematician's AI Review Experiment: 97.7% of 266 Auto-Generated Comments Flag Real Issues — littmath · 2026-08-27
- AI audit identifies 97.7% real issues in math papers — littmath · 2026-08-27
- [source] AI audit not infallible: Refine missed a known lemma error in paper — littmath · 2026-08-27