96% of LLM verdicts repeat Jev's most confident errors, experiments show

deliprao · x · 2026-09-26

Delip Rao's key experimental finding: on Jev's most confident errors, 96.0% of LLM verdicts repeat its wrong answer, versus about 50% if errors were independent. The number is the most direct quantification of his argument that LLM judges make highly correlated errors, undermining cascade and ensemble setups built around the new decision model.

Related event: Experiments Show LLM Judge Errors Are Highly Correlated, Limiting Cascades(3 posts)→

Original post →

More from Research

Research channel →