LangChain's Jev evaluator cuts agent eval score variance by up to 913x at 1/80th the cost
LangChain · x · 2026-09-20
- LangChain introduces Jev-as-a-Judge, an agent eval method that returns typed answers instead of generating text like LLM judges
- On continuous scoring, Jev's variance was 92–913x lower than GPT-5.6 Luna, Terra, and Claude Sonnet 4.6
- It averaged 0.44s and $0.00035/call — $0.34 total vs $28.17 for Claude
- Authors caution results are early but say it points to a compelling new direction for agent evals
More from coding & agent
- Dev picks C for new engine: not for AI, but full control and universal bindings — gdechichi · 2026-09-20
- Dev take: AI's million-lines-a-day code is slop, public perception lags — BLUECOW009 · 2026-09-20
- chatpipe-mcp Lets AI Coding Agents Publish Live Pages With Shareable URLs — modelcontextprotocol · 2026-09-20
- Agentic Web talk: generations already replace Chrome with AI assistants — EdenEmarco177 · 2026-09-20
- A Mac drawing app added an MCP server to test its tools—and it became the killer feature — DrawSimple_for_MacOS · 2026-09-20
- Hugo Bowne offers free 30-minute Lightning Lesson on AI agent evals — hugobowne · 2026-09-20