Anthropic adds official eval-building and hillclimbing workflow to Claude Code skill
ClaudeDevs · x · 2026-09-29
Anthropic (claude.dev blog, by Lance Martin) added two commands to the claude-api skill so Claude Code can build evaluations and improve apps against them automatically:
- /claude-api build-eval: set up an evaluation inside your codebase
- /claude-api hillclimb: improve the app one change at a time, with a held-out example set to catch overfitting
Key eval design principles from the article:
- Eval tasks must mirror production — not be chosen because they're easy to generate or grade
- Stronger models should score higher; if not, tasks are ambiguous or the grader is miscalibrated
- Keep passable headroom at the frontier — the best model at max effort should sit well below 100%, or you can't reliably judge the impact of changes
An official, hands-on guide to eval engineering and agent iteration workflows.
Related event: Anthropic Enables Automated Eval Building with Claude Code(2 posts)→
More from coding & agent
- Bret Taylor's Sierra moves enterprise agents into Slack with proactive Ghostwriter — ZackRW · 2026-09-29
- Dev after 2 days: Claude is excellent, Codex great for long-horizon tasks but poorly designed — cneuralnetwork · 2026-09-29
- TradingView MCP server hits 4.7k stars, wiring Claude and Cursor into live market data and backtesting — tom_doerr · 2026-09-29
- Dev uses Opus 5.5 to build a 3rd-person MOBA mixing LoL, DOTA and HoTS heroes — TAbrodi · 2026-09-29
- He Built a Claude Code Skill That Acts as an AI Automation Consultant — First Run Showed 63,000% ROI — ricoesrico · 2026-09-29
- No App Needed: 15MB Brain Connectome Model Docked to a Poster with Plain JavaScript — Dr_Alex_Crimi · 2026-09-29