How to Build Better AI Evals with Claude Code in 5 Steps

petergyang · x · 2026-08-23

This episode features Shreya Shankar and Hamel Husain, experts who have taught AI evals to over 4,500 engineers. They discuss the fundamentals of AI evaluation, emphasizing that while core principles remain (looking at real data), AI agents now assist in thoughtful analysis. The podcast covers a live audit of evals, top-down vs. bottom-up strategies, and demos a free skill for running evals in Claude Code.

Related event: Experts discuss building AI evaluations with Claude Code(2 posts)→

Original post →

More from coding & agent

coding & agent channel →