AI audit of 19 papers: 97.7% of comments identified real issues
littmath · x · 2026-08-27
Daniel Litt published the overall results of an automated review of his papers by Refine. Across 19 works, the AI generated 266 detailed comments: 239 were correct, 21 partially correct, and 6 incorrect, meaning 97.7% identified a real issue. Most findings were local errors, with only 2 theorem-preserving defects and 5 technical corrections to main results found. No fundamental failures were detected.
More from Research
- Best practices for reliable critic training in LLM RL — heghbalz · 2026-08-27
- Tiny 307M-parameter model outperforms 26x larger Qwen in embedding benchmarks — lateinteraction · 2026-08-27
- Study: Generative AI shifts UK students' major choices — RishiBommasani · 2026-08-27
- EvoMax Evolves Compact Genome Editor to 97% Efficiency Using Sparse Data — bravo_abad · 2026-08-27
- Bixbench3: Frontier Agents Score Below 50% in Reproducing Paper Analysis — xeophon · 2026-08-27
- Navigator n2 released: 27B model achieves frontier computer use performance — DhruvBatra_ · 2026-08-27