SciSlopBench Flags AI-Written Papers at 85.9% Accuracy, Correlates With Lower ICLR Scores

SeoulNatlUniv · hf · 2026-10-05

Seoul National University built SciSlopBench (390 AI-generated papers, each paired with a human paper matched by problem and contribution type), measuring "scientific slop" via six measures across Structure, Argument, and Artifacts: each part looks plausible while the connecting scientific reasoning breaks down.

Findings

Mitigation: direct metric optimization fails (standard revisions leave residual slop; slop-aware prompting triggers reward hacking). SciSlopHarness guides a fixed LLM to revise only where experiment records support changes, cutting the AI-human gap by 63% without human reference targets.

Original post →

More from Safety

Safety channel →