Experiments show LLM judge cascades gain at most 2 points as errors correlate

deliprao · x · 2026-09-26

Delip Rao reports experimental results: even with tuned confidence thresholds, cascading gains at most 1.5 points over the best single judge with cross-fitted thresholds, and at most 2.0 points with oracle thresholds. Cascading to juries (judge ensembles) doesn't help much either, because all the models make correlated errors.

Related event: Experiments Show LLM Judge Errors Are Highly Correlated, Limiting Cascades(3 posts)→

Original post →

More from Research

Research channel →