Using Jev as an LLM judge fails: near-uniform high scores diverge from humans

A developer found Jev unusable as an automated scorer on OntBench: it gave nearly all outputs high scores, contradicting both human ratings and Codex's scores.

2026-09-18 ~ 2026-09-18 · 3 related posts