Experiments show LLM judge cascades gain at most 2 points as errors correlate
deliprao · x · 2026-09-26
Delip Rao reports experimental results: even with tuned confidence thresholds, cascading gains at most 1.5 points over the best single judge with cross-fitted thresholds, and at most 2.0 points with oracle thresholds. Cascading to juries (judge ensembles) doesn't help much either, because all the models make correlated errors.
Related event: Experiments Show LLM Judge Errors Are Highly Correlated, Limiting Cascades(3 posts)→
More from Research
- Redwood Research: Astra reasons far better with filler tokens, outside its chain-of-thought — scaling01 · 2026-09-26
- ACuRL: zero-human-data continual learning for computer-use agents lands at NeurIPS — ysu_nlp · 2026-09-26
- Mathematician Tivadar Danka shares 10 biggest lessons from 20 years in mathematics — TivadarDanka · 2026-09-26
- GPT-6 Luna uses fewer reasoning tokens than 5.6 on ARC-AGI-2, hard tasks stymie both — mhmazur · 2026-09-26
- Contrastive World Models: latent-space world models without pixel prediction — bonniesjli · 2026-09-26
- Researchers including Google build first complete brain map of a male fruit fly, 166,000+ neurons — burny_tech · 2026-09-26