LangChain team: evals are the #1 blocker keeping companies from becoming AI-native
hwchase17 · x · 2026-10-06
LangChain's Brace Sproul says evals are the biggest blocker in most agent development pipelines, with changes coming to LangSmith "in a few short weeks."
An accompanying post lays out the core definitions: an eval is a grader applied to a trace; a grader is anything that scores agent performance (from a simple code check to an LLM-as-judge); a trace is the full record of one agent run—input, every tool call (e.g. searchorders, issuerefund), and output.
The takeaway: hand-tweaking prompts and testing a few inputs only gets you so far—evals are what get agents deployment-ready and keep them working in production, and they're the top hurdle to companies becoming truly AI-native.
Related event: LangChain Team Highlights Evals as Top Barrier to Enterprise AI Adoption(2 posts)→
More from coding & agent
- Teknium catalogs 305 open, hackable devices your AI agent can control and live in — Teknium · 2026-10-06
- ESR: LLMs make closed-source binaries 'transparent as glass', ending secrecy as we know it — nitarshan · 2026-10-06
- Crawler + GPT agent scanned 5,137 jobs for $6.53, found 3 low-competition niches — Arindam_1729 · 2026-10-06
- xAI Cookbook adds 4 Grok SDK examples: screenshot-to-React, X sentiment tracker, AI ad generator — tetsuoai · 2026-10-06
- Power user runs a swarm of personal agents — poke, instinct, muse, dot, Grok bot, pickle and more — sebkrier · 2026-10-06
- omni-macos: A Fully Local Semantic Finder for Apple Silicon, Now Daily-Driver Good — JinaAI_ · 2026-10-06