Agent Evals: Jev Is 180x Cheaper Than LLM Judges, Only 2.6 Points Less Accurate

ialijr · reddit · 2026-10-10

The author replayed 193 real eval verdicts on Jev and three LLM judges (Sonnet, Haiku, Gemini Flash Lite), scored against hand labels, 5 runs each:

A cascade of Jev first + Sonnet on the uncertain 16% achieves Sonnet-level accuracy at 1/6 the cost. Verdict: not a replacement, but an excellent first pass. The report includes full methodology, per-rubric-type results, calibration and cascade charts, a code example for porting your own rubrics, and an open repo to reproduce every number.

Original post →

More from coding & agent

coding & agent channel →