LangSmith ships Jev-as-a-judge to score every production agent trace at low cost
LangChain · x · 2026-09-22
LangChain announced Jev-as-a-judge in LangSmith Evals, a low-cost way to evaluate open-ended agent behavior at scale. Key points: score every production trace instead of sampling; check more criteria per trace without cost climbing; catch safety/security issues fast enough to trigger automated responses.
The post reviews the history of agent evals — code-based checks (fast but only cover pre-specified behavior) vs LLM-as-a-judge (flexible but slower, pricier, non-deterministic) — and positions Jev as a "System One" model judge balancing speed, cost and coverage, with setup steps via the Evaluators tab.
Related event: LangChain Introduces Jev-as-a-Judge for Cheap, Fast Agent Evals(9 posts)→
More from coding & agent
- Building an image rating tool with GPT Vision and Jev: what worked and what didn't — huangyun_122 · 2026-09-22
- Dev ditches throttled GitHub Actions, open-sources dsr for local releases via act — doodlestein · 2026-09-22
- Nous ships official plugin to run Hermes Agent on Claude Pro/Max subscriptions, no API key — Teknium · 2026-09-22
- Kos raises $12M from 8VC to build AI finance agents for data center invoices — ahelkky · 2026-09-22
- Harrison Chase on decision models: agentic systems are just good engineering around models — Hacubu · 2026-09-22
- LangSmith Integrates Jev for Cheap Large-Scale Trace Mining That Feeds Directly Into Evals — hwchase17 · 2026-09-22