Decision grader replaces LLM judge: 32x cheaper, 8x faster, 94% agreement on evals
rhythmrg · x · 2026-10-06
Applied Compute's RL and eval pipelines previously graded with an LLM judge producing one paragraph per rubric criterion — slow and pricey for complex tasks.
They swapped in a decision grader on a code QA benchmark in AC2: one call yielding yes/no probabilities for every criterion at once. Results on 100 sample tasks (407 criteria):
- 32x cheaper and 8x faster
- 94% agreement with the original LLM judge; 99% when the decision grader was confident
- Low-confidence (30-70%) items flagged ambiguous rubric entries, letting the team identify noisy tasks and clean the dataset
Takeaway: a cheap, fast, calibrated grader lets you grade more rollouts and ask more questions per criterion than per-item LLM judging.
More from coding & agent
- Garry Tan ports Doom to Paul Graham's Bel LISP in 20 minutes using Opus 5.5 — garrytan · 2026-10-07
- L0pht Hacker Chris Wysopal on Securing AI-Written Code: 'Make It Secure' Isn't a Prompt — WeldPond · 2026-10-07
- Designing evals for contract review agents: ten lawyers, ten redlines — graceisford · 2026-10-07
- Engineer reviews 20 agent-written PRs a day: 'I didn't sign up to be a full-time proofreader' — Specialist_Agent3599 · 2026-10-07
- Researcher unveils RSI paradigm: generic disposable agents plus an evolving knowledge base — yisongyue · 2026-10-07
- KNOWS Benchmark Jointly Evaluates Agents on Search, Tools and Visual Understanding — anmarasovic · 2026-10-07