Decision grader replaces LLM judge: 32x cheaper, 8x faster, 94% agreement on evals

rhythmrg · x · 2026-10-06

Applied Compute's RL and eval pipelines previously graded with an LLM judge producing one paragraph per rubric criterion — slow and pricey for complex tasks.

They swapped in a decision grader on a code QA benchmark in AC2: one call yielding yes/no probabilities for every criterion at once. Results on 100 sample tasks (407 criteria):

Takeaway: a cheap, fast, calibrated grader lets you grade more rollouts and ask more questions per criterion than per-item LLM judging.

Original post →

More from coding & agent

coding & agent channel →