Jev as Judge: Structured Decisions Replace Text Generation for Agent Evals, With Open-Source Demo

blaizedsouza · x · 2026-09-25

A clear walkthrough of Jev vs LLM-as-Judge: both need the same evidence (request, policy, tool results, agent answer), but an LLM judge generates verdicts token by token while Jev returns bounded structured decisions. The companion open-source repo patchy631/jev-as-judge replays ten frozen support traces—including invented policies, false action claims, and evaluator-injection attempts—runs deterministic checks before paying for semantic evaluation, and logs results as auditable Comet Opik experiments.

Related event: Jev-as-a-Judge Cuts Agent Evaluation Cost With Decision-Based Judging(3 posts)→

Original post →

More from coding & agent

coding & agent channel →