LLM Judge Bias Inflates False Positives

IanArawjo · x · 2026-07-19

The author warns: **do not blindly trust unverified LLM judge results**. Simulations across various data distributions revealed that biased judges cause a massive spike in false positives (indicated by the tall bars). Without correction, seemingly "significant" findings are likely false. The accompanying chart displays false positive rates across different settings: a human oracle stays near the nominal 0.05 significance level, whereas several LLM-based approaches show error rates far exceeding 0.05, sometimes approaching 0.9 or 1.0.

Related event: Study Warns of High False Positive Rates in Unchecked LLM Judges(7 posts)→

Original post →

More from Research

Research channel →