CMU Paper: Decision-Only LLM Judge Matches SOTA at 0.36% of the Cost via Cascade
CarnegieMellonU · hf · 2026-09-23
A CMU paper studies JEV-as-a-Judge: a decision-only judge as an economical first pass for LLM evaluation. Against 16 generative and reward-model judges with blinded human adjudication, it stays within 3 points of a SOTA LLM judge on preference and evidence-grounded factuality tasks at 0.36% of the cost. Gaps concentrate on low-confidence decisions and derivation-checking; a frozen cascade that accepts confident verdicts and escalates uncertain ones retains 99% accuracy at lower cost.
Related event: Jev-as-a-Judge cuts evaluation cost 43% with 1% accuracy loss(2 posts)→
More from Research
- ECCV 2026 paper studies which high-dimensional latents suit diffusion models — _akhaliq · 2026-09-25
- LeWAM: Lightweight World Action Model Hits 92.28% Success on RoboTwin 2.0 — udmrzn · 2026-09-25
- SmolDataEnvs open-sources 5K+ verifiable RL tasks for training small code models — ben_burtenshaw · 2026-09-25
- ROTATE Ad Hoc Teamwork Training Paper Wins NeurIPS 2026 Spotlight — cuijiaxun · 2026-09-25
- Open-source Jive rebuilds the agentic loop around "System One" graph calls, beating Codex on speed and cost 7x — burakyildizdev · 2026-09-25
- Complexity theorists breach an 'invisible fence' with Human-AI collaboration in new PRG paper — fortnow · 2026-09-25