Hugo Bowne offers free 30-minute Lightning Lesson on AI agent evals
hugobowne · x · 2026-09-20
Hugo Bowne-Anderson is running a free 30-minute Lightning Lesson, "AI Agent Evals: Test What Matters for Your Agent," following up on his podcast with Hamel Husain and a guest post on the Vanishing Gradients blog.
Topics include:
- Turning tasks and failures into evals, and when to use code checks, LLM judges, or human review
- What changes when evaluating coding, research, conversational, or computer-use agents
- Using traces and evals to decide what to change in your harness and verify the improvement
Sessions: Sept 21, 3pm Pacific / Sept 22, 8am Sydney, with live attendance or a recording.
More from coding & agent
- LangChain's Jev-as-a-Judge: a cheaper, more precise alternative to LLM judges for agent evals — LangChain · 2026-09-20
- Fintech engineer asks: where does sensitive production data go when AI agents plug in? — noexz · 2026-09-20
- jevcache memoizes model decisions by (model, schema, state) — repeat calls cost $0 and return in ~0ms — JiliJeanlouis · 2026-09-20
- GitHub quietly adds Copilot HydraFusion preview, and one blog post about it drew 59% of the week's traffic — shashib · 2026-09-20
- Dev laments no agent harness has built-in budgeting, even a simple /budget command — arthurcolle · 2026-09-20
- Dev swaps jev into real browser-use harness: it breaks instantly on real websites — TheZachMueller · 2026-09-20