Study Proposes Preregistration Protocol to Curb LLM Evaluation P-Hacking

As LLMs are widely used as annotation tools or evaluation judges, p-hacking (data-mining style statistical manipulation) in scientific research has become increasingly severe, and the proliferation of end-to-end autonomous AI scientists further amplifies this risk. To address this issue, Maria Thomas from JHU and Nihar Shah from CMU released a study proposing a preregistration template and anti-p-hacking protocol designed specifically for LLM-based research, aiming to promote stricter scientific methods and improve reproducibility.

Protocol Mechanism and Key Details

The core protocol requires researchers to preregister their analysis plan and a set of "eligible models," then run confirmation experiments on the first eligible LLM released after preregistration. Because the model does not exist at the time of commitment, it mechanically prevents targeted overfitting or evaluation manipulation. Additionally, the team released a preregistration template tailored for LLM-based research to prevent post-hoc interpretation bias and dataset leakage.

Efficacy and Cost

On two tasks with "zero true effects" (where any significant result is a p-hack), the protocol was tested across 20 models, 4 providers, and 11 configurations, successfully blocking 73.9% and 72.7% of p-hacking migration. The research team's own preregistered experiments showed that 85.7% of configurations previously p-hacked on older models were successfully blocked when using the next eligible model. The protocol retains the ability to detect true effects, with its main cost being the wait for the next model release, which has a median waiting time of about 10 days.

2026-07-06 ~ 2026-07-06 · 7 related posts