Jev-as-a-Judge: cheap judge cascade keeps 99% of GPT-6 accuracy at 57% of the cost
dair_ai · x · 2026-09-24
dairai highlights the Jev-as-a-Judge paper, whose core finding is that you should use a cheap judge for most evals and escalate only uncertain calls to a frontier model.
The paper measures JEV, TypeSafe AI's decision-only judge, in a cascade:
- On 510 held-out preference pairs, a cascade accepting JEV's confident verdicts and escalating the rest to GPT-6 Astra retained 99% of GPT-6's accuracy at about 57% of its fee.
- JEV returns verdicts and label probabilities with no reasoning text: $0.044 per 1,000 judgments at 0.152s median latency, versus $12.182 and 1.885s for GPT-6 — about 277x cheaper.
- On ordinary preference and evidence-grounded factuality it stays within 3 points of GPT-6 (92.2% vs 93.5% on RewardBench, 87.5% vs 86.7% on HaluEval), but the gap grows to 9-20 points on tasks requiring derivation checks.
Related event: Jev-as-a-Judge cuts evaluation cost 43% with 1% accuracy loss(2 posts)→
More from Models
- LiquidAI extends lossless speculative decoding to vision-language models — JosephJacks_ · 2026-09-25
- GPT-5.2 solves a COLT 2022 open problem the researcher had chased since 2016 — kfountou · 2026-09-25
- Agents bypass monitoring guardrails with strategies that improve as reasoning effort scales — maksym_andr · 2026-09-25
- micro1 launches flow-transform 1.0, hits 96.0% F1 on PrivacyBench PII transformation — omarsar0 · 2026-09-25
- UkisAI ships Swift reasoning LLM family: -63.4% thinking tokens at 1.8x speed on Qwen base — Secure_Recording_472 · 2026-09-25
- "How Many Versions of Humanity's Last Exam Before They Rename It?" — code_star · 2026-09-25