Preregistered Analysis and Eligible Model Sets Prevent Eval p-Hacking

krisgligoric · x · 2026-07-06

A new protocol has been proposed to prevent p-hacking in LLM evaluations: preregister an analysis plan alongside a set of "eligible models," then run confirmation experiments on the first eligible LLM released after preregistration. Because the model did not exist at the time of commitment, it cannot be specifically gamed, mechanically eliminating the manipulation of evaluations.

Related event: Study Proposes Preregistration Protocol to Curb LLM Evaluation P-Hacking(7 posts)→

Original post →

More from Research

Research channel →