96% of LLM verdicts repeat Jev's most confident errors, experiments show
deliprao · x · 2026-09-26
Delip Rao's key experimental finding: on Jev's most confident errors, 96.0% of LLM verdicts repeat its wrong answer, versus about 50% if errors were independent. The number is the most direct quantification of his argument that LLM judges make highly correlated errors, undermining cascade and ensemble setups built around the new decision model.
Related event: Experiments Show LLM Judge Errors Are Highly Correlated, Limiting Cascades(3 posts)→
More from Research
- Redwood Research: Astra reasons far better with filler tokens, outside its chain-of-thought — scaling01 · 2026-09-26
- ACuRL: zero-human-data continual learning for computer-use agents lands at NeurIPS — ysu_nlp · 2026-09-26
- Mathematician Tivadar Danka shares 10 biggest lessons from 20 years in mathematics — TivadarDanka · 2026-09-26
- GPT-6 Luna uses fewer reasoning tokens than 5.6 on ARC-AGI-2, hard tasks stymie both — mhmazur · 2026-09-26
- Contrastive World Models: latent-space world models without pixel prediction — bonniesjli · 2026-09-26
- Researchers including Google build first complete brain map of a male fruit fly, 166,000+ neurons — burny_tech · 2026-09-26