LangChain's Jev-as-a-Judge: cheap, fast semantic verifiers for agent evals
multiply_matrix · x · 2026-09-20
LangChain published "Jev-as-a-Judge for Agent Evals" by Daniel Shea and Seán Roche, amplified by hwchase17.
- Jev is a fundamentally different kind of evaluator: it returns typed answers directly instead of generating text like an LLM judge.
- The team benchmarked Jev against LLM judges on accuracy, repeatability, latency, and cost.
- Best fit is online evals, where grading lots of traces cheaply and quickly matters more than verbose LLM-judge reasoning.
Related event: LangChain Introduces Jev-as-a-Judge for Agent Evals(5 posts)→
More from coding & agent
- Dev Uses Computer-Use Agent to Install/Uninstall 300+ Mac Apps, Adds 90 Cleanup Rules — vista8 · 2026-09-20
- Spotify ships Xirp, a Mac app born from 36,000 AI coding sessions — shashib · 2026-09-20
- Resetwatch plugin aggregates usage limits and reset times for 14 AI providers in one page — Teknium · 2026-09-20
- An inbox re-ranked live by importance instead of time, built with jev — xkonjin · 2026-09-20
- DAIR.AI launches auto-updating Awesome Jev collection, curated by Jev itself — omarsar0 · 2026-09-20
- Open-source Tampermonkey scripts add paste-to-upload for WeChat, Xiaohongshu and Douyin — vista8 · 2026-09-20