LangChain's Chase RTs guide: evals are the #1 blocker to AI-native companies
hwchase17 · x · 2026-10-06
In a long thread RT'd by LangChain founder hwchase17, businessbarista argues evals are the #1 thing keeping companies from becoming AI-native: tweaking prompts and hand-testing a few inputs only goes so far — evals get agents deployment-ready and keep them working in production.
Core definitions:
- Eval: a grader applied to a trace
- Grader: anything scoring agent performance, from plain code checks to LLM-as-judge
- Trace: full record of one agent run — input, every step, output (e.g. refund lookup → issuerefund → reply)
The thread walks through grader families, starting with deterministic checks (pure code, no model), noting real setups typically stack several graders on the same trace. An accompanying guide on eval-driven development is linked.
Related event: LangChain Team Highlights Evals as Top Barrier to Enterprise AI Adoption(2 posts)→
More from coding & agent
- DeepSeek-style agent infrastructure ran 1.3B sandboxes in 4 weeks, peaking at 170K concurrent — teortaxesTex · 2026-10-06
- DIY agentic memory with an nginx proxy: surviving 35 compactions — cryowastakenbycryo · 2026-10-06
- Burkov: no one will need to fix AI code — just prompt it to fix itself — burkov · 2026-10-06
- "A single Claude Code session won't replace your software": a reality check on the one-shot AI narrative — seatedro · 2026-10-06
- Your evals are your product spec: the most common mistake AI product teams make — realmadhuguru · 2026-10-06
- Dev's Ultrafast review: 6x token cost, $500 plan hits weekly limits every 1-2 days — holdenmatt · 2026-10-06