LangSmith ships Jev-as-a-judge to score every production agent trace at low cost

LangChain · x · 2026-09-22

LangChain announced Jev-as-a-judge in LangSmith Evals, a low-cost way to evaluate open-ended agent behavior at scale. Key points: score every production trace instead of sampling; check more criteria per trace without cost climbing; catch safety/security issues fast enough to trigger automated responses.

The post reviews the history of agent evals — code-based checks (fast but only cover pre-specified behavior) vs LLM-as-a-judge (flexible but slower, pricier, non-deterministic) — and positions Jev as a "System One" model judge balancing speed, cost and coverage, with setup steps via the Evaluators tab.

Related event: LangChain Introduces Jev-as-a-Judge for Cheap, Fast Agent Evals(9 posts)→

Original post →

More from coding & agent

coding & agent channel →