PPI-corrected tests can beat human-only labels, but only with enough judge alignment
IanArawjo · x · 2026-07-27
The post explains that the figure plots the statistical power of four PPI-corrected tests in the EvalStats framework, compared with their classical versions using only human labels. The key message is that PPI can detect real effects by combining human and LLM judgments while avoiding an inflated false-positive risk.
The table shows that the efficiency gain depends on label alignment and sample size. Gains are largest when the human-only sample is small and judge agreement is high (around 70%–80%), but the method can still help at moderate alignment around 50%. In contrast, when alignment is low, the multiplier approaches 1x, meaning there is little or no benefit over classical testing.
Related event: LLM Judges Require Over 50% Human Alignment for Statistical Efficacy(3 posts)→
More from Research
- A 1951 mechanical tortoise is being used to explain today’s LLM scaling walls — mtizard · 2026-07-27
- Cheap storage makes SCD Type 2 look obsolete, says a Meta-style data engineer — Zachly · 2026-07-27
- Creed-Bench launches as a new eval for personal context — craighepburn · 2026-07-27
- Reddit asks whether continued pretraining, SFT or RL works best on Qwen3.6-27B — No-Paper-557 · 2026-07-27
- Researchers observe models often need an a→b→c→d path before discovering a simple arithmetic trick — dejavucoder · 2026-07-27
- Claude Code is not reliable enough for long research projects without human supervision — _akpiper · 2026-07-27