Hamel Husain open-sources evals Skills that let coding agents build product-specific evals
sh_reya · x · 2026-09-23
Lenny Rachitsky spotlighted evals-skills, open-sourced by Hamel Husain and Shreya Shankar: a set of skills that guide AI coding agents to build product-specific evals, distilled from helping 50+ companies. The evals-start entry skill routes to eval-audit (inspect existing pipelines) or error-discovery (build annotation interfaces and sample traces). Payoffs cited: Ramp lifted receipt-collection accuracy 35%→83%; Shopify's AI workflow builder is 2.2x faster and 68% cheaper; Harvey nearly doubled its internal quality score; Cursor cut Auto Balance costs 41%. Lenny also notes half of 25 PM job openings now ask for evals experience.
Related event: Hamel Husain Open-Sources evals-skills for Coding Agents(2 posts)→
More from coding & agent
- Okta launches Blueprint Alliance with AWS, CrowdStrike, Wiz to secure AI agents — yenkel · 2026-09-23
- LangChain Demos Shopping Agent Using Stagehand and Stripe Link CLI — LangChain · 2026-09-23
- Cognition SWE-2 fixes a bug by deleting all its code — and that's the right answer — silasalberti · 2026-09-23
- Codex drives a VCV Rack synth live over MCP in an AI-performed patch demo — gabrielchua · 2026-09-23
- DevFest talk asks: do websites need an agent-first paradigm as AI agents browse and buy — philnash · 2026-09-23
- Developer shares his note-taking workflow: one daily inbox note, capture first, organize later — dSebastien · 2026-09-23