Experiments Show LLM Judge Errors Are Highly Correlated, Limiting Cascades

Delip Rao's experiments show cascading LLM judges yields at most 1.5 points over the best single judge, with 96% of judges repeating the same confident errors—echoing ICML 2025 findings that errors across 350+ models are highly correlated.

2026-09-26 ~ 2026-09-26 · 3 related posts