jevals: locally runnable agent evals using JEV-style decision models

byebaybay · reddit · 2026-09-24

A new open-source project, jevals, proposes replacing expensive LLM-as-a-judge agent evaluations with locally runnable JEV-style decision models (one-shot classifiers). The author argues generation-based judging is too costly and slow at millions of decisions, that natural-language-adaptable classifiers generalize across tasks without per-task fine-tuning, and that classification/regression remains the right tool when the answer is a label, probability, or score.

Related event: Ex-Apple Engineer Open-Sources jevals for Local Agent Evaluation(2 posts)→

Original post →

More from coding & agent

coding & agent channel →