Using Jev as an LLM judge fails: near-uniform high scores diverge from humans
A developer found Jev unusable as an automated scorer on OntBench: it gave nearly all outputs high scores, contradicting both human ratings and Codex's scores.
2026-09-18 ~ 2026-09-18 · 3 related posts
- Jev fails as an LLM scorer on OntBench: rates almost everything positively, contradicting human and Codex ratings — amplifiedamp · 2026-09-18
- Jev as an LLM judge flops: scores nearly everything positively, disagrees with humans — amplifiedamp · 2026-09-18
- Using an LLM as benchmark scorer fails: over-optimistic ratings diverge from human judgment — amplifiedamp · 2026-09-18