Why LLM judge ensembles amplify bias: correlated models boost systemic errors

IanArawjo · x · 2026-09-07

IanArawjo hypothesizes why ensembling many LLM judges backfires: their biases positively correlate due to model correlation, amplifying a latent systemic bias. At high inter-rater reliability there's less noise, giving that bias more room to express itself — producing bogus super-significant results.

Related event: High Consistency Across LLM Judge Ensembles Can Amplify Bias, Study Warns(2 posts)→

Original post →

More from Research

Research channel →