Google Cloud's Agent Clinic: build an automated eval suite for a LangGraph agent in 60 min
LangChain · x · 2026-10-01
In Episode 3 of Google Cloud Tech's Agent Clinic, Dani Zamora and mattferoz demonstrate building an automated eval suite for a docs agent built with LangChain/LangGraph, with Merge API Gateway as the intelligence provider.
Key point: terminal test runs won't catch multi-turn agent regressions. The video lays out a 4-step framework for going from vibes to a benchmark for any AI agent. The author also ran the agent on three OSS coding tools: t3dotcodes, opencode, and pidotdev.
More from coding & agent
- All the Copyright Easter Eggs in That AI Music Video Came From Claude — technollama · 2026-10-01
- ChatGPT Wrote Lyrics, Suno Made Music, Claude Coded a 5,200-Frame Music Video — technollama · 2026-10-01
- New tool makes benchmarking across factory configs trivial, not just models — vikvang1 · 2026-10-01
- Magnitude inference engine hits #1 on HN, claims up to 2x faster local open-model runs than llama.cpp — nickbaumann_ · 2026-10-01
- marimo-lens lets you point at charts and steer your coding agent directly — S_Conradi · 2026-10-01
- Non-coder builds MCP-only medieval trade game played entirely by AI agents — Wainfare · 2026-10-01