Study Proposes Preregistration Protocol to Curb LLM Evaluation P-Hacking
As LLMs are widely used as annotation tools or evaluation judges, p-hacking (data-mining style statistical manipulation) in scientific research has become increasingly severe, and the proliferation of end-to-end autonomous AI scientists further amplifies this risk. To address this issue, Maria Thomas from JHU and Nihar Shah from CMU released a study proposing a preregistration template and anti-p-hacking protocol designed specifically for LLM-based research, aiming to promote stricter scientific methods and improve reproducibility.
Protocol Mechanism and Key Details
The core protocol requires researchers to preregister their analysis plan and a set of "eligible models," then run confirmation experiments on the first eligible LLM released after preregistration. Because the model does not exist at the time of commitment, it mechanically prevents targeted overfitting or evaluation manipulation. Additionally, the team released a preregistration template tailored for LLM-based research to prevent post-hoc interpretation bias and dataset leakage.
Efficacy and Cost
On two tasks with "zero true effects" (where any significant result is a p-hack), the protocol was tested across 20 models, 4 providers, and 11 configurations, successfully blocking 73.9% and 72.7% of p-hacking migration. The research team's own preregistered experiments showed that 85.7% of configurations previously p-hacked on older models were successfully blocked when using the next eligible model. The protocol retains the ability to detect true effects, with its main cost being the wait for the next model release, which has a median waiting time of about 10 days.
2026-07-06 ~ 2026-07-06 · 7 related posts
- LLM 让科研 p-hacking 更严重 — krisgligoric · 2026-07-06
- [source] 提出预注册分析+合格模型集防评测p-hack — krisgligoric · 2026-07-06
- [source] 该协议在两类任务阻断p-hack迁移超72% — krisgligoric · 2026-07-06
- 自测预注册实验85.7%曾p-hack配置被阻止 — krisgligoric · 2026-07-06
- 防p-hack协议保留检测力但需等约10天 — krisgligoric · 2026-07-06
- 自主AI科学家加剧p-hacking风险 — krisgligoric · 2026-07-06
- [source] 研究发布LLM实验预注册模板,推动科学可重复性标准 — krisgligoric · 2026-07-06