Hamel Husain: Jev works for Evals — an LLM judge is just a classifier

HamelHusain · x · 2026-09-20

Hamel Husain answers whether Jev can be used for evals: yes — an LLM judge is fundamentally a classifier, so validate it against human labels and don't overfit. He links his FAQ on why binary (pass/fail) evaluations beat 1-5 Likert scales: adjacent rating differences are subjective and inconsistent across annotators, statistical detection needs larger samples, and annotators default to midpoints. To track gradual progress, split sub-metrics into separate binary checks (e.g. "4 of 5 expected facts included") instead of using a scale.

Related event: LangChain Launches Jev-as-a-Judge for Agent Evals, Now Live in LangSmith(9 posts)→

Original post →

More from coding & agent

coding & agent channel →