FabScore Evaluates Fabrication in AI Papers
muhao_chen · x · 2026-07-11
FabScore for Evaluating Fabrication in AI Papers
This post introduces FabScore: a metric and method for fine-grained evaluation of "fabrication" in automated AI research. Accepted as a spotlight paper at the ICML 2026 AI for Science workshop and selected as one of the best papers, with the author noting an upcoming morning session presentation at COEX in Seoul.
Key Findings
- Evaluated 144 AI-generated papers from sources including Sakana AI's AI Scientist, MLR-Bench, Analemma AI's FARS, and the 2025 Agents4Science Open Conference.
- In 54 real conference submissions, about 70% contained at least one fabrication.
- Even among accepted papers, this proportion reached 59.3%.
Links to the paper and code are also provided.
More from Research
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11
- VidMap uses RoMa coarse matching on all frames, fine-scale only for keyframes — ducha_aiki · 2026-09-11
- Bug Hunt Bench author: leaderboard noise is about 2-3 points — PawelHuryn · 2026-09-11
- PNAS paper shows a tiny billiard-ball system is a universal computer — undecidability lives in two dimensions — eigensteve · 2026-09-11
- New paper: Absolute pose estimation from affine cues and gravity direction — ducha_aiki · 2026-09-11