CMU paper: small model judges LLM outputs at 0.36% of the strongest judge's cost
akshay_pachaar · x · 2026-09-27
A new Carnegie Mellon paper tests whether small models like Jev can handle bounded evaluation decisions — grounding checks, instruction following, response preference — instead of full LLM-as-judge setups.
Key findings:
- Close on focused judgments: Jev stayed within 3 percentage points of the strongest judge tested on ordinary response preferences, evidence-grounded factuality, and final-answer checks.
- Massive cost gap: on the paper's matched workload, Jev cost only 0.36% of the strongest comparator's fee.
The practical takeaway: use cheap judges to evaluate far more agent responses across more criteria, escalating only the hard cases to stronger models.
More from Research
- FuseReg: Layer-Fusion Regularization Cuts gFID up to 29% in Representation Autoencoders — USC-PSI-Lab · 2026-09-28
- Kaggle Game Arena: Evaluating LLMs via Head-to-Head Chess, Poker, and Werewolf — kaggle · 2026-09-28
- InternW0-Δ: A World Action Model Trained on 20K+ Hours of Open Robot Data, Fully Open-Sourced — Xingyu Miao · 2026-09-28
- Berkeley's Morphometric Imitation Hits 89.3% Zero-Shot Real-World Success Across 3 Robot Hands — Berkeley · 2026-09-28
- Fermi estimate: brain may pack 500k-5M molecular switching units per 'parameter' — JosephJacks_ · 2026-09-28
- Contrastive World Models: swapping pixel reconstruction for InfoMax boosts robustness — burny_tech · 2026-09-28