LangChain's 38-minute masterclass on building agent evals and environments
LangChain · x · 2026-09-15
LangChain amplified a 38-minute evals masterclass from its Labs lead, breaking down why evals now drive AI progress from frontier labs to open source.
Key points:
- An eval = did the agent do the job? Built from Tasks (checkable jobs) and Verifiers (scripts, models, or humans with checklists)
- Modern agents do far more autonomous work than 3 years ago, requiring Environments that faithfully capture real work
- Good evals remain hard to build; frontier companies pay millions to acquire them across domains
- Every team that wants to own their intelligence needs to own their evals to measure and improve agents over time
Related event: LangChain Labs Lead Gives 38-Minute Masterclass on Agent Evals(2 posts)→
More from coding & agent
- Running 12 market-making bots side by side: clock-based vs distance-based oracle refresh on BTC-BRL — cardosofede · 2026-09-15
- Four-Dev Team Shares Hard Lessons on Collaborating with AI Coding Agents in Real Time — Common-Pizza3787 · 2026-09-15
- How Do You Detect an Agent That's Alive but Cooked in Production? — ousco · 2026-09-15
- You're overpaying for Claude Code; Quotient claims 35-50% token savings — ycombinator · 2026-09-15
- Turnstone launches free AI workspace with persistent local agents — ycombinator · 2026-09-15
- OpenAI says 2 engineers + Codex rewrote habitat service in Rust, 6x CPU and 15x memory gains — imjustnewatai · 2026-09-15