FabScore Evaluates Fabrication in AI Papers
muhao_chen · x · 2026-07-11
FabScore for Evaluating Fabrication in AI Papers
This post introduces FabScore: a metric and method for fine-grained evaluation of "fabrication" in automated AI research. Accepted as a spotlight paper at the ICML 2026 AI for Science workshop and selected as one of the best papers, with the author noting an upcoming morning session presentation at COEX in Seoul.
Key Findings
- Evaluated 144 AI-generated papers from sources including Sakana AI's AI Scientist, MLR-Bench, Analemma AI's FARS, and the 2025 Agents4Science Open Conference.
- In 54 real conference submissions, about 70% contained at least one fabrication.
- Even among accepted papers, this proportion reached 59.3%.
Links to the paper and code are also provided.
More from Research
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11