Anthropic adds build-eval and hillclimb commands to automate eval design with Claude

RLanceMartin · x · 2026-09-29

Anthropic's Lance Martin published a guide on designing evals and hillclimbing without fooling yourself. The claude-api skill now offers /claude-api build-eval to create evals inside your codebase and /claude-api hillclimb to improve your app one change at a time, with held-out examples to catch overfitting. Key principles: tasks should mirror production; stronger models and more thinking should score higher (if not, tasks or graders are miscalibrated); and the frontier model at max effort should sit well below 100% so improvements remain measurable.

Related event: Anthropic Enables Automated Eval Building with Claude Code(2 posts)→

Original post →

More from coding & agent

coding & agent channel →