Swapping the LLM Judge From Sonnet to a 16x Cheaper Model: Zero Verdict Flips Across 84 Gradings
EvalRaccoonDev · reddit · 2026-09-28
An engineer working on Coder Eval (Apache 2.0) replaced their Sonnet-based LLM judge ($1.65 per run) with GPT Luna plus a short calibration prompt ($0.09) — replaying 84 real gradings produced zero verdict flips. The bigger finding: 15 of 85 gradings used scores the rubric never defines (0.9/0.95 on a 1.0/0.8/0.5/0.2/0 scale), and 10 required a perfect score to pass. Every flip, on every judge model, traced back to those threshold bugs. Recommendation: fix your scoring rubric before shopping for a judge — a cheap one may suffice.
More from coding & agent
- Motif: an MCP design library that lets coding agents consult design guidance while they build UI — Able_Bus_5988 · 2026-09-28
- Forge tracks how AI agents discover and buy your x402 APIs — kleffew94 · 2026-09-28
- Open-source Aseprite MCP server lets agents animate directly on a live canvas — 72-frame rain scene stress test — Complex-Log-635 · 2026-09-28
- Cloudflare's Kitesurf Agentic Browser Adds WebMCP Support and Terminal Rendering — dinasaur_404 · 2026-09-28
- OpenAI President Brockman's prompting guide reworked for the agentic era — kimmonismus · 2026-09-28
- Open-source qiaomu-cut Skill turns one sentence into a reproducible AI video production pipeline — vista8 · 2026-09-28