AI audit not infallible: Refine missed a known lemma error in paper
littmath · x · 2026-08-27
Evaluating the completeness of the Refine reports, Daniel Litt noted they are definitely not perfect. The tool missed an error in a lemma in one of his published papers, previously found by Jordan Ellenberg and Alex Smith (and subject to an erratum, though it didn't affect a main result). This indicates that while AI audits are highly effective, human review is still necessary to catch missed nuances.
More from Research
- Best practices for reliable critic training in LLM RL — heghbalz · 2026-08-27
- Tiny 307M-parameter model outperforms 26x larger Qwen in embedding benchmarks — lateinteraction · 2026-08-27
- Study: Generative AI shifts UK students' major choices — RishiBommasani · 2026-08-27
- EvoMax Evolves Compact Genome Editor to 97% Efficiency Using Sparse Data — bravo_abad · 2026-08-27
- Bixbench3: Frontier Agents Score Below 50% in Reproducing Paper Analysis — xeophon · 2026-08-27
- Navigator n2 released: 27B model achieves frontier computer use performance — DhruvBatra_ · 2026-08-27