Swapping the LLM Judge From Sonnet to a 16x Cheaper Model: Zero Verdict Flips Across 84 Gradings

EvalRaccoonDev · reddit · 2026-09-28

An engineer working on Coder Eval (Apache 2.0) replaced their Sonnet-based LLM judge ($1.65 per run) with GPT Luna plus a short calibration prompt ($0.09) — replaying 84 real gradings produced zero verdict flips. The bigger finding: 15 of 85 gradings used scores the rubric never defines (0.9/0.95 on a 1.0/0.8/0.5/0.2/0 scale), and 10 required a perfect score to pass. Every flip, on every judge model, traced back to those threshold bugs. Recommendation: fix your scoring rubric before shopping for a judge — a cheap one may suffice.

Original post →

More from coding & agent

coding & agent channel →