Mathematician's AI Review Experiment: 97.7% of 266 Auto-Generated Comments Flag Real Issues
littmath · x · 2026-08-27
Mathematician Daniel Litt published the full record of running Refine.ink on 19 of his papers: of 266 detailed comments, 239 were correct, 21 partially correct and only 6 wrong — 97.7% identified a real issue. A ChatGPT audit plus manual spot-checks turned these into 216 numbered erratum items. He calls the output "slop" but useful, noting AI often over-edits (adding a paragraph where one word would do).
By impact: 209 were local correctable errors (statement, proof step, formula, citation, hypothesis); the audit found 2 substantial theorem-preserving defects and 5 technical corrections to main results, with no fundamental failures. He adds that most errors came from badly propagated edits while optimizing, and that one known lemma error found by human peers (Jordan Ellenberg, Alex Smith) was missed by Refine. More errata await coauthor permission.
More from Research
- IJCAI 2026 Tutorial: How LLMs are reshaping mathematical optimization — StanfordAILab · 2026-08-27
- Research predicts optimal model size and data allocation for pre-training — yoavgo · 2026-08-27
- ARC-AGI-3 learnings: Key designs for self-learning AI agents — GregKamradt · 2026-08-27
- Recommended Reads: Poggio on Minds & Machines, Welling on GenAI Thermodynamics — neurovium · 2026-08-27
- Category Theory Master: AI Expands Mathematical Truth but Doesn't Change Core — begusgasper · 2026-08-27
- Tempus Cancer Model Boosts Survival Prediction AUC with 1.67M Patient Dataset — neuroecology · 2026-08-27