CMU paper: small model judges LLM outputs at 0.36% of the strongest judge's cost

akshay_pachaar · x · 2026-09-27

A new Carnegie Mellon paper tests whether small models like Jev can handle bounded evaluation decisions — grounding checks, instruction following, response preference — instead of full LLM-as-judge setups.

Key findings:

The practical takeaway: use cheap judges to evaluate far more agent responses across more criteria, escalating only the hard cases to stronger models.

Original post →

More from Research

Research channel →