LangSmith ships Jev-as-a-judge to score every production trace cheaply
airesearch12 · x · 2026-09-22
LangChain launched Jev-as-a-judge in LangSmith: score every production trace instead of a sample, check more criteria per trace without rising costs, and catch safety or security issues fast enough to trigger automated responses. Harrison Chase says it scores traces cheaply and accurately. Jev is a zero-shot classifier with frontier-ish intelligence.
Related event: LangSmith Launches Jev-as-a-judge for Cheap Full-Trace Evaluation(2 posts)→
More from coding & agent
- SemIf open-sources Jev-style semantic ifs: 4B model runs on a single 3090, browser demo live — Hacubu · 2026-09-22
- Developer shares model-split workflow: Perplexity for research, Claude Code as the coding workhorse — ZabihullahAtal · 2026-09-22
- Fireworks shows two Jev training recipes: GRPO-style reward scoring and offline DPO/SFT filtering — sophiamyang · 2026-09-22
- After a year of building, the end-state agent harness: max-freedom execution backend plus a free-form canvas UI — TheZachMueller · 2026-09-22
- Grok 4.7 hits 140K installs as open-source IDE extensions; AFK Pilot rides the wave — PawelHuryn · 2026-09-22
- After a year of AST-RAG papers, Chonks indexes a whole codebase into one SQLite file — _solidude · 2026-09-22