Agent Evals: Jev Is 180x Cheaper Than LLM Judges, Only 2.6 Points Less Accurate
ialijr · reddit · 2026-10-10
The author replayed 193 real eval verdicts on Jev and three LLM judges (Sonnet, Haiku, Gemini Flash Lite), scored against hand labels, 5 runs each:
- Jev: 90.6% accuracy at $0.07 per 1k verdicts
- Sonnet 4.6: 93.2% at $13.21 per 1k (180x pricier)
- Jev made 0 errors when confident (125 of 191 items)
A cascade of Jev first + Sonnet on the uncertain 16% achieves Sonnet-level accuracy at 1/6 the cost. Verdict: not a replacement, but an excellent first pass. The report includes full methodology, per-rubric-type results, calibration and cascade charts, a code example for porting your own rubrics, and an open repo to reproduce every number.
More from coding & agent
- Debugging 9K lines of 100% AI-generated code: standard tactics failed to find root cause — blaizedsouza · 2026-10-11
- ManimGX open-sources a Rust+wgpu 3D animation engine for agents, 90× faster than ManimCE — Scobleizer · 2026-10-11
- Redot Engine welcomes AI-generated PRs — but only if contributors understand their own code — esrtweet · 2026-10-11
- $500 ex-mining BC-250 cluster runs Qwen 35B at 145 tok/s with 256k context — Ok-Breadfruit-3523 · 2026-10-11
- Jev founder's harness guide: coding agents up to 200x faster, 400x cheaper — blaizedsouza · 2026-10-11
- First 24 hours with Hark agent: auto-connects accounts, fixes its own Notion errors, 2FA remains a hurdle — Scobleizer · 2026-10-11