JEV-as-a-Judge: confidence-routed cascade beats GPT-6 accuracy at 41% of the cost
kalyan_kpl · x · 2026-10-11
A new arXiv paper from a CMU team proposes JEV-as-a-Judge: a decision-only judge that returns label probabilities instead of text, using its confidence to decide whether to accept its verdict or escalate to a reasoning judge.
- Against 16 generative and reward-model judges with blinded human adjudication, JEV comes within 3 points of GPT-6 wherever a verdict can be read off the text — at 0.36% of its fee and a 0.15s median latency
- It falls behind on tasks requiring derivation (math, code, logic); its confidence marks this boundary
- With a pre-frozen threshold, accepting confident verdicts and escalating the rest is 0.9 points more accurate than GPT-6 on 1,610 held-out pairs at 41% of the fee, and a pre-specified live test on two new workloads matched GPT-6's accuracy exactly
- Caveats: confidence routing weakens on style-adversarial pairs and reference-free prose; the paper closes with a simple recipe for validating thresholds locally
Related event: JEV-as-a-Judge Nears GPT-6 Accuracy at a Fraction of Cost(2 posts)→
More from Models
- Musk shows Grok researching and ordering Lego Star Wars kits in one prompt — elonmusk · 2026-10-11
- Google ships EmbeddingGemma 2: 740M multimodal embeddings that run on phones — dl_weekly · 2026-10-11
- Rumor: Grok 4.8 with 2.5T parameters (up 67% from Grok 4.6) may launch this week — mark_k · 2026-10-11
- TensorFold 1.0.7 writes each learned fact into ~10 new neurons, 4x faster with 3D view — HankYeomans · 2026-10-11
- Arena scores look close: OpenAI 88 vs Claude 83 means double the error rate — i_dg23 · 2026-10-11
- Rumor: Anthropic's next-gen Fable 5.5 can one-shot SVG generation — koltregaskes · 2026-10-11