Benchmarking Claude Code /goal: judge said done 17/17, 8 were broken

Sanechka_SS · reddit · 2026-10-07

The author examined Claude Code's /goal command: a small model (Haiku) reads only the transcript each turn to decide if the goal is met — it never runs commands or opens files, so a confident 'tests pass' from the agent is most of its evidence.

Benchmark: 4 small coding tasks with hidden checks, Sonnet 5.5 and Haiku 4.5, every run graded 3 times. The judge said 'met' in 17 of 17 plain runs; 8 were broken (1 Sonnet, 7 Haiku), the worst passing only 4% of hidden checks.

Fix — goalpost, an MIT-licensed hooks plugin: the goal becomes criteria with a command each; stopping is blocked until every criterion has a passing check run after the last edit; existing tests are protected from quiet edits; a fresh-context auditor re-runs everything at the end (striking pair: 28/34 vs 34/34). Results: Sonnet went from 8/9 fully correct to 9/9 (within noise); Haiku went 79.8%→89.7% on hidden checks but stayed unreliable, partly because its auditor is also Haiku. Roughly doubles time and cost.

Lesson: an LLM judge is only as good as the evidence you put in front of it — 'the agent said so' isn't evidence. Plugin and full run dataset are open source.

Original post →

More from coding & agent

coding & agent channel →