build-eval + hillclimb bring real eval tooling to Claude Code agents

Arindam_1729 · x · 2026-09-29

A practical eval tooling combo for Claude Code: build-eval auto-generates evals from production traces, while hillclimb iterates to push accuracy up and cost down. It accompanies Anthropic's 45-minute "Evals for Taste" session, which walks through wiring a slide-generation Managed Agent, scoring it against SlidesBench, and iterating prompts from failures — turning "this looks bad" into a number you can optimize.

Related event: Anthropic Adds Automated Eval Building and Optimization Commands to Claude Code(4 posts)→

Original post →

More from coding & agent

coding & agent channel →