Free 30-min lesson details: choosing between code checks, LLM judges and human review for agent evals
hugobowne · x · 2026-09-21
Hugo Bowne-Anderson shared registration details for his free 30-minute Lightning Lesson, "AI Agent Evals: Test What Matters for Your Agent" (Sep 21/22):
- Turn real tasks and failures into evals using code checks, LLM judges, and human review
- Match evals to agent types — coding, research, conversational, computer-use — and see where generic scores fall short
- Inspect traces, identify harness changes, and rerun evals to verify fixes without breaking existing tasks
The instructor is data scientist Hugo Bowne-Anderson (Vanishing Gradients host, has taught engineers from Netflix, Meta, Amazon); 204 students enrolled so far.
Related event: Agent eval success rates hinge on how you define success(2 posts)→
More from coding & agent
- Delta and OpenTable block AI agents, hinting at a coming platform stand-off — Scobleizer · 2026-09-21
- ForgePoint Signal launches MCP-native daily US estate and tax law monitoring — modelcontextprotocol · 2026-09-21
- Memory MCP adds persistent semantic memory to AI assistants via Turso and OpenAI — modelcontextprotocol · 2026-09-21
- Gemini CLI PR fixes queued tool calls still executing after scheduler disposal — andreivince · 2026-09-21
- GPT-6-controlled robot arm breaks gripper during LEGO grasping; team seeks force-aware control — paigeinsf · 2026-09-21
- Agent failures don't throw: a 7-year engineer on clean double-runs and zero accountability — ilien-dev · 2026-09-21