Hamel Husain open-sources evals-skills to let coding agents audit and build your AI evals
RichmanRonald · x · 2026-09-23
Hamel Husain and Shreya Shankar released evals-skills (835 GitHub stars), a set of Skills that guide AI coding agents to build product-specific AI evals, distilled from helping 50+ companies and teaching thousands of students.
- evals-start is the entry point, routing you to the right skill
- eval-audit inspects an existing eval pipeline and catches common footguns—authors recommend at least running an audit; low-hanging fruit is common
- error-discovery builds a customized annotation interface and helps sample traces intelligently for error analysis
- Installable as plugins for Claude and Codex
Lenny Rachitsky endorsed it as saving 'many hours and a lot of mistakes.'
Related event: Hamel Husain Open-Sources evals-skills for Coding Agents(2 posts)→
More from coding & agent
- AI engineering is more like lawmaking than board games, argues Drew Breunig — dbreunig · 2026-09-23
- Seroter's daily digest: GPT-6 and Opus 5.5 ship, 1 in 4 agents run unmonitored — rseroter · 2026-09-23
- Spawning Claude agents that auto-open terminal panes: 'tmux can't do this' — letandrewcook · 2026-09-23
- Toddler's interactive storybook built in two hours with an agent workflow — mimi10v3 · 2026-09-23
- Your data stack is about to get less forgiving: agents need data that's true now — bigdata · 2026-09-23
- Opinion: Agents make software good at using software, not just being used — r0ck3t23 · 2026-09-23