Mathematician Daniel Litt Audits 19 of His Own Papers with AI, Finds 97.7% of Comments Flag Real Issues
数学家 Daniel Litt 于 8 月 27 日公开了一项完整实验:用 AI 工具 Refine.ink 配合 ChatGPT 自动审查自己 19 篇论文并生成勘误。结果显示 266 条详细评论中 239 条完全正确、21 条部分正确、仅 6 条错误,即 97.7% 指出了真实问题,其中约 5 处可视为影响主要结果的严重问题。
已确认
- 审查覆盖 19 篇论文,产出 266 条评论:239 条正确、21 条部分正确、6 条错误,识别真实问题比例 97.7%。
- 5 处严重问题多为缺失技术性假设(如主定理缺少射影性或约化性假设),Litt 认为细心读者本可推断出这些隐含假设,但确实必须修正。
- 大多数发现属于轻微问题:错字、歧义表述等。
- Litt 还用 AI 生成了半自动勘误(errata),他坦承这些勘误是「slop」、仅经轻度人工编辑,抽查显示大体正确但往往过于激进(如整段替换而非补一个词),目前保持原样发布,以展示几小时努力能达到的程度。
- 关于错误成因,Litt 复盘称其首篇长论文中几乎全部错误源于优化结果时未把编辑正确传播到全文各处,是典型的论文出错方式。
- Refine 并不完整:它遗漏了 Litt 一篇已发表论文引理中的错误,该错误此前已被 Jordan Ellenberg 和 Alex Smith 发现并发布勘误,未影响主要结果。
为什么重要
- 这是一次数学家对自身全部论文的系统化 AI 审计,提供了量化数据(97.7% 有效率),显示 AI 审稿在发现真实技术缺陷上已具实用价值。
- 实验同时暴露局限:AI 会遗漏已知错误、生成的勘误过于激进,人工核查仍不可省略。
2026-08-27 ~ 2026-08-27 · 8 related posts
Primary sources
- [source] AI audit flags 5 serious math issues, incl. missing theorem hypotheses — littmath · 2026-08-27
- Mathematician Litt: Paper Errors Mostly From Badly Propagated Edits, Not Deep Flaws — littmath · 2026-08-27
- Mathematician's AI-generated errata are "slop" but mostly correct — littmath · 2026-08-27
- A few hours of effort yields AI-generated errata for a paper archive — littmath · 2026-08-27
- [source] AI audit of 19 papers: 97.7% of comments identified real issues — littmath · 2026-08-27
- Mathematician's AI Review Experiment: 97.7% of 266 Auto-Generated Comments Flag Real Issues — littmath · 2026-08-27
- AI audit identifies 97.7% real issues in math papers — littmath · 2026-08-27
- [source] AI audit not infallible: Refine missed a known lemma error in paper — littmath · 2026-08-27