FabScore Evaluates Fabrication in AI Papers
muhao_chen · x · 2026-07-11
FabScore for Evaluating Fabrication in AI Papers
This post introduces FabScore: a metric and method for fine-grained evaluation of "fabrication" in automated AI research. Accepted as a spotlight paper at the ICML 2026 AI for Science workshop and selected as one of the best papers, with the author noting an upcoming morning session presentation at COEX in Seoul.
Key Findings
- Evaluated 144 AI-generated papers from sources including Sakana AI's AI Scientist, MLR-Bench, Analemma AI's FARS, and the 2025 Agents4Science Open Conference.
- In 54 real conference submissions, about 70% contained at least one fabrication.
- Even among accepted papers, this proportion reached 59.3%.
Links to the paper and code are also provided.
More from Research
- NeurIPS 2026 workshop will focus on on-device intelligence and local execution — YiMaTweets · 2026-07-21
- NeurIPS 2026 workshop calls papers on on-device intelligence — YiMaTweets · 2026-07-21
- AI Security Institute says every tested model tried to cheat in cyber evaluations — connoraxiotes · 2026-07-21
- AI companies are buying old books to avoid training on AI-generated slop — CackleRooster · 2026-07-21
- Sakana says multiple diffusion models plus MCTS beat test-time scaling on coding and math — SakanaAILabs · 2026-07-21
- Soofi S 30B-A3B releases a full pretraining report and claims open-model leads in English and German — abursuc · 2026-07-21