Preregistered Analysis and Eligible Model Sets Prevent Eval p-Hacking
krisgligoric · x · 2026-07-06
A new protocol has been proposed to prevent p-hacking in LLM evaluations: preregister an analysis plan alongside a set of "eligible models," then run confirmation experiments on the first eligible LLM released after preregistration. Because the model did not exist at the time of commitment, it cannot be specifically gamed, mechanically eliminating the manipulation of evaluations.
Related event: Study Proposes Preregistration Protocol to Curb LLM Evaluation P-Hacking(7 posts)→
More from Research
- MaP-WAM tackles non-Markovian robot manipulation with memory-grounded planning — Sizhe Zhao · 2026-09-11
- Negative Self-Distillation improves LLM reasoning by avoiding flawed reasoning paths — Rongcan Pei · 2026-09-11
- DeepMind-led paper makes design docs the source of truth, code disposable — SMART regenerates in 1.5-3h for ~$100 — Roger_M_Taylor · 2026-09-11
- GameWorld wins Best Paper Runner-Up at ECCV 2026 Multimodal Digital Agents Workshop — MikeShou1 · 2026-09-11
- Yann LeCun live at ECCV on World Models — Weak_Assistance_5261 · 2026-09-11
- 3D ResNet Paper Crosses 3,000 Citations Eight Years After CVPR 2018 — HirokatuKataoka · 2026-09-11