CMU paper: cheap judge cascade keeps 99% accuracy at 0.36% of the cost
Stefania_druga · x · 2026-09-25
A CMU paper, "JEV-as-a-Judge: Accept When Confident, Escalate When Unsure," tackles LLM-as-a-judge costs at scale:
- Compared against 16 generative and reward-model judges with blinded human adjudication, the decision-only JEV judge lands within 3 percentage points of the strongest LLM judge on ordinary preference and evidence-grounded factuality tasks — at just 0.36% of its fee.
- Gaps emerge on tasks requiring derivation-checking or resisting well-written wrong answers.
- Crucially, JEV's errors concentrate in low-confidence decisions. A frozen cascade — accept confident verdicts, escalate uncertain ones to a stronger LLM — retains 99% of the strong judge's accuracy at a fraction of the cost.
The sharer expects this "cheap first-pass judge + escalation" routing pattern to become very common. Results are workload-specific, not a universal guarantee.
More from Infra
- Anthropic signs $11.6B Akamai cloud deal, compute commitments top $517B in 11 months — The Decoder · 2026-09-25
- Not every task needs frontier models: local Qwen 4 27B is pulling users away — haider1 · 2026-09-25
- Oracle on the hook to pay data centre investors even if site has no electricity — Betelbuddy · 2026-09-25
- Under $2K for the GPU part: MCIO cables to retimers keep full x16 per card — TheZachMueller · 2026-09-25
- Oracle Hedges Against Stargate Data Center Delays, Exposing AI Infra Bottleneck — TansuYegen · 2026-09-25
- Anthropic Commits $11.6B to Akamai for Cloud Infra, Deal May Reach $20B — TansuYegen · 2026-09-25