Jev as an LLM judge flops: scores nearly everything positively, disagrees with humans

amplifiedamp · x · 2026-09-18

A developer tried using Jev as a scorer in OntBench, a benchmark for how well different LLMs build ontology maps, and found it doesn't work: it scores almost everything very positively and disagrees with both the author's manual ratings and Codex's ratings (all scorers blinded).\n\nThe author shared the full scoring prompt — including evidence boundary instructions and a five-level rubric — asking the community whether something is wrong with the setup.

Related event: Using Jev as an LLM judge fails: near-uniform high scores diverge from humans(3 posts)→

Original post →

More from Models

Models channel →