Eval engineering cuts grading costs to $77.81, but judges still disagree with themselves 13.6% of the time

BLUECOW009 · x · 2026-07-29

The thread summarizes an article on eval engineering for agents: two researchers reportedly replaced $7,500 of human grading with only $77.81 in model calls, but found that production judges are still unreliable.

Key findings:

The post argues that agent teams should treat evaluation as an engineering layer, not a one-off script:

The broader point: evals become a core part of the agent system you already pay for.

Original post →

More from coding & agent

coding & agent channel →