Mathematician's AI Review Experiment: 97.7% of 266 Auto-Generated Comments Flag Real Issues

littmath · x · 2026-08-27

Mathematician Daniel Litt published the full record of running Refine.ink on 19 of his papers: of 266 detailed comments, 239 were correct, 21 partially correct and only 6 wrong — 97.7% identified a real issue. A ChatGPT audit plus manual spot-checks turned these into 216 numbered erratum items. He calls the output "slop" but useful, noting AI often over-edits (adding a paragraph where one word would do).

By impact: 209 were local correctable errors (statement, proof step, formula, citation, hypothesis); the audit found 2 substantial theorem-preserving defects and 5 technical corrections to main results, with no fundamental failures. He adds that most errors came from badly propagated edits while optimizing, and that one known lemma error found by human peers (Jordan Ellenberg, Alex Smith) was missed by Refine. More errata await coauthor permission.

Related event: Mathematician Daniel Litt Audits 19 of His Own Papers with AI, Finds 97.7% of Comments Flag Real Issues(8 posts)→

Original post →

More from Research

Research channel →