Hamel Husain ships eval skills: run /eval-audit to find low-hanging fruit in your pipeline
HamelHusain · x · 2026-09-24
AI engineers Hamel Husain and shreya released a set of skills for building and auditing LLM evals, recommending everyone at least run an /eval-audit on their existing pipeline—they consistently find low-hanging fruit this way. randalolson notes evals are "no longer a skill issue" thanks to this.
More from coding & agent
- Garry Tan: coding harnesses aren't syntactic sugar — fixes land faster with no added in-editor time — garrytan · 2026-09-24
- Simon Willison: GPT-6 Astra built 5K/10K running routes in 27 min, but compaction hid the code — dl_weekly · 2026-09-24
- Designer uses Opus 5.5 to rebuild his site's hero shader with cursor-stirred smoke and HDR particles — AIandDesign · 2026-09-24
- FOSS battlemap-mcp lets Claude Code and Codex drive Dungeondraft map editor — thekannenG · 2026-09-24
- Garry Tan says capy makes his PR workflow 4x faster than raw Codex/Claude Code — garrytan · 2026-09-24
- Read-only MCP servers for Proxmox and pfSense: agents hit the API, not screenshots — Mustela__ · 2026-09-24