jevals: locally runnable agent evals using JEV-style decision models
byebaybay · reddit · 2026-09-24
A new open-source project, jevals, proposes replacing expensive LLM-as-a-judge agent evaluations with locally runnable JEV-style decision models (one-shot classifiers). The author argues generation-based judging is too costly and slow at millions of decisions, that natural-language-adaptable classifiers generalize across tasks without per-task fine-tuning, and that classification/regression remains the right tool when the answer is a label, probability, or score.
Related event: Ex-Apple Engineer Open-Sources jevals for Local Agent Evaluation(2 posts)→
More from coding & agent
- Reading an invoice and moving money should not share one permission: agent auth principles — TechNadu · 2026-09-24
- ComfyUI Launches Comfy Router: One API for Image, Video, 3D and Audio Models — cpaik · 2026-09-24
- Agent ignored the GitHub plugin for computer use — fixed by adding a rule to AGENTS.md — jdjohnson · 2026-09-24
- Hamel Husain's AI Evals FAQ: model benchmarks and product evals answer different questions — HamelHusain · 2026-09-24
- Open-source MCP server gives agents page-change tracking with SHA-256 capture certificates — Einperegrin · 2026-09-24
- LangChain Leans Into 'Composable AI': Jev and System One Models in the Agent Harness — hwchase17 · 2026-09-24