build-eval + hillclimb bring real eval tooling to Claude Code agents
Arindam_1729 · x · 2026-09-29
A practical eval tooling combo for Claude Code: build-eval auto-generates evals from production traces, while hillclimb iterates to push accuracy up and cost down. It accompanies Anthropic's 45-minute "Evals for Taste" session, which walks through wiring a slide-generation Managed Agent, scoring it against SlidesBench, and iterating prompts from failures — turning "this looks bad" into a number you can optimize.
More from coding & agent
- 10 Open Source Repos That Give Your AI Agent the Whole Internet — victor_explore · 2026-09-30
- 3 Applied AI Pro Tips: Explore With Frontier Models, Scale With Small Ones — brandon_galang · 2026-09-30
- First-principles thinking remains the human edge in the agentic coding era — rseroter · 2026-09-30
- Mollick: OpenAI's DevDay agent vision was obsolete within a year — emollick · 2026-09-30
- Karpathy Coins 'Vibe Coding': Give In to the Vibes, Forget the Code Exists — alexmacgregor__ · 2026-09-30
- PSClaudeCode: an open-source Claude Code clone in 250 lines of PowerShell — dfinke · 2026-09-30