Benchmarking Claude Code /goal: judge said done 17/17, 8 were broken
Sanechka_SS · reddit · 2026-10-07
The author examined Claude Code's /goal command: a small model (Haiku) reads only the transcript each turn to decide if the goal is met — it never runs commands or opens files, so a confident 'tests pass' from the agent is most of its evidence.
Benchmark: 4 small coding tasks with hidden checks, Sonnet 5.5 and Haiku 4.5, every run graded 3 times. The judge said 'met' in 17 of 17 plain runs; 8 were broken (1 Sonnet, 7 Haiku), the worst passing only 4% of hidden checks.
Fix — goalpost, an MIT-licensed hooks plugin: the goal becomes criteria with a command each; stopping is blocked until every criterion has a passing check run after the last edit; existing tests are protected from quiet edits; a fresh-context auditor re-runs everything at the end (striking pair: 28/34 vs 34/34). Results: Sonnet went from 8/9 fully correct to 9/9 (within noise); Haiku went 79.8%→89.7% on hidden checks but stayed unreliable, partly because its auditor is also Haiku. Roughly doubles time and cost.
Lesson: an LLM judge is only as good as the evidence you put in front of it — 'the agent said so' isn't evidence. Plugin and full run dataset are open source.
More from coding & agent
- Caveman Renders AI Agent Skills as Images, Cutting Token Cost by 61% — bigaiguy · 2026-10-07
- SOTA models quietly nerfed ~15 times in 2025, dev suggests dual Codex+Claude subscriptions — StewartalsopIII · 2026-10-07
- Kev: an open-source reimplementation of Jev built on Qwen3.5 that runs 100% locally — Arindam_1729 · 2026-10-07
- Google Releases SAM, a P2P Network Letting AI Agents Discover and Call Each Other's Tools — thisguyknowsai · 2026-10-07
- STEER: A human-in-the-loop rewrite desk that fixes AI slop from your inline marks — bfrench · 2026-10-07
- Manufacturers say skip the new agents: process mining plus RPA saved 12,000 hours a year — shashib · 2026-10-07