Jev as Judge: Structured Decisions Replace Text Generation for Agent Evals, With Open-Source Demo
blaizedsouza · x · 2026-09-25
A clear walkthrough of Jev vs LLM-as-Judge: both need the same evidence (request, policy, tool results, agent answer), but an LLM judge generates verdicts token by token while Jev returns bounded structured decisions. The companion open-source repo patchy631/jev-as-judge replays ten frozen support traces—including invented policies, false action claims, and evaluator-injection attempts—runs deterministic checks before paying for semantic evaluation, and logs results as auditable Comet Opik experiments.
Related event: Jev-as-a-Judge Cuts Agent Evaluation Cost With Decision-Based Judging(3 posts)→
More from coding & agent
- Block joins x402 Foundation, contributes Lightning payments for agentic commerce — kleffew94 · 2026-09-25
- Stripe engineer answers 'what's it like at Stripe' with demos over memos culture — schwentker · 2026-09-25
- Wasmer runs a real PostgreSQL 18.4 server on iOS and in the browser via WebAssembly — jedisct1 · 2026-09-25
- Google Cloud details 4 patterns for real-time AI voice agents that see, talk, think and code — rseroter · 2026-09-25
- Opus 5.5 coded a polished 30-second hype video from one prompt in 3 hours of autonomous work — DaveRogenmoser · 2026-09-25
- Letting an agent redesign git: cosm makes the code graph, not files, the source of truth — rseroter · 2026-09-25