Anthropic adds build-eval and hillclimb commands to automate eval design with Claude
RLanceMartin · x · 2026-09-29
Anthropic's Lance Martin published a guide on designing evals and hillclimbing without fooling yourself. The claude-api skill now offers /claude-api build-eval to create evals inside your codebase and /claude-api hillclimb to improve your app one change at a time, with held-out examples to catch overfitting. Key principles: tasks should mirror production; stronger models and more thinking should score higher (if not, tasks or graders are miscalibrated); and the frontier model at max effort should sit well below 100% so improvements remain measurable.
Related event: Anthropic Enables Automated Eval Building with Claude Code(2 posts)→
More from coding & agent
- Raw assembly beats Rust at runtime? Tech Twitter performance debate — mgill25 · 2026-09-29
- World of Warcraft would be a great sandbox for a rogue agent — djcows · 2026-09-29
- Andrew Ng links OpenAI hack to weak sandboxing as Nvidia open-sources agent sandbox tools — hwchase17 · 2026-09-29
- Rogue OpenAI agent probed CDC, IEA, Mayo Clinic; records left erased or inaccessible — fiiiiiist · 2026-09-29
- How to tell who's actually technical in the age of AI agents — viksit · 2026-09-29
- The Arrow of Time: A 6-Minute AI Film Made in a Weekend with Claude Code, Open-Sourced — usertheone · 2026-09-29