Paper: Classifier Judge Matches LLM Judges and Copies 96% of Their Errors
deliprao · x · 2026-09-28
A new paper, "JEV vs. LLMs as Rubric Judges: Cheaper, Faster, and Wrong in the Same Places," introduces JEV, an evaluator that skips text generation entirely and outputs probabilities directly over a fixed answer set, processing all criteria in a single request.
Key findings:
- JEV matches commercial LLM judges on binary decisions while being cheaper and faster, but graded scoring exposes a ceiling.
- When JEV makes a confident error, expensive LLM judges reproduce the exact same mistake 96% of the time.
- Graded accuracy deficits stem from shared scale compression, and automated judges systematically grade lower than humans.
- Highly correlated errors severely limit the value of fallback judge cascades; scaling evaluator compute is useless when benchmarks rely on unwritten criteria.
More from Models
- GPT-6 Sol appears on LMArena: 24-hour Direct Mode window before Battle and Agent Mode — arena · 2026-09-28
- Kaggle Game Arena: Google's LLM benchmark pits models against each other in chess, poker, werewolf — weballergy · 2026-09-28
- Qwen3.8-27B goes live on Nebius Token Factory for agent workflows — HowDevelop · 2026-09-28
- AISI: GPT-6 Astra ran unsanctioned supply-chain attacks in simulated cyber evals — ShakeelHashim · 2026-09-28
- Codex Computer Use 'Neutered' by Guardrails; Opus 5.5 Does the Job on First Try — iannuttall · 2026-09-28
- Shanghai AI Lab ships Intern-Decision multimodal family, 4B model beats Jev 1.13.0 — stingning · 2026-09-28