Kitaru Workshop Walks Through Full Agent Eval Loop: From Real Traces to Verified Fixes
strickvl · x · 2026-08-27
The Kitaru team hosted a live workshop, "Debug, Replay, Improve," using a customer-support agent to demonstrate a complete agent evaluation loop:
- Import and explore real agent traces
- Identify a specific behavior worth improving
- Turn the finding into an evaluator
- Update the agent and replay the same examples
- Compare experiment results to verify the change actually worked
Around 67 people registered. No prior Kitaru experience was required, and it targeted developers building agents.
More from coding & agent
- Alibaba Open-Sources CommerceAgentBench for Long-Horizon Agent Testing — mhdfaran · 2026-08-27
- CommerceAgentBench Challenges Agents to Process 300 Messy Procurement Emails — mhdfaran · 2026-08-27
- Docker Is Not a Real Sandbox for Agent Code: From Containers to microVMs — aidenclarke_12 · 2026-08-27
- Warmwind Demo: AI Agents Interacting via Screen Without APIs — Med1_Ai · 2026-08-27
- Developer Builds Road Trip Simulator with Claude, Animates Routes in Real Time — vinishkapoor · 2026-08-27
- H3 Max tested: 15s 768p clips in ~30s, fastest video model right now — techhalla · 2026-08-27