PPI-corrected tests can beat human-only labels, but only with enough judge alignment

IanArawjo · x · 2026-07-27

The post explains that the figure plots the statistical power of four PPI-corrected tests in the EvalStats framework, compared with their classical versions using only human labels. The key message is that PPI can detect real effects by combining human and LLM judgments while avoiding an inflated false-positive risk.

The table shows that the efficiency gain depends on label alignment and sample size. Gains are largest when the human-only sample is small and judge agreement is high (around 70%–80%), but the method can still help at moderate alignment around 50%. In contrast, when alignment is low, the multiplier approaches 1x, meaning there is little or no benefit over classical testing.

Related event: LLM Judges Require Over 50% Human Alignment for Statistical Efficacy(3 posts)→

Original post →

More from Research

Research channel →