Full Clay × LangChain conversation on evaluating millions of monthly agent runs
LangChain · x · 2026-09-10
LangChain published the full conversation with Clay: the company runs millions of agent runs monthly, where manual review breaks down, so it uses LangSmith's online evaluators and is testing the insights product to understand agent behavior. Jeff Barg walks through Clay's agent evaluation workflow.
More from coding & agent
- Mollick: you can't be in the loop for long-running agents, but you can oversee it — emollick · 2026-09-10
- Ethan Mollick on steering long-running agents: when to instruct, queue or fork — emollick · 2026-09-10
- The mental model for LLM guardrails: a separate layer that distrusts the model — Careless_Sabfey_4906 · 2026-09-10
- Ethan Mollick: you're probably not steering your long-running coding agents enough — emollick · 2026-09-10
- AgentGrad: intervention-guided prompt optimization hits SOTA with 2.5x faster tuning — _akhaliq · 2026-09-10
- Claude embeds cyber controls in models vs Codex as a separate orchestratable model — HankYeomans · 2026-09-10