Jev fails as an LLM scorer on OntBench: rates almost everything positively, contradicting human and Codex ratings

amplifiedamp · x · 2026-09-18

A developer tried using Jev as an automated scorer in OntBench (a benchmark for how well different LLMs build ontology maps) and found it unusable: it scores almost everything very positively, disagreeing with both the author's manual ratings and Codex's ratings.

The full five-level scoring prompt (from 'multiple major mistakes' to 'supported claims with explicit limits') is shared, asking the community what went wrong. It's a textbook LLM-as-judge failure: the scorer lacks discrimination and skews positive.

Related event: Using Jev as an LLM judge fails: near-uniform high scores diverge from humans(3 posts)→

Original post →

More from Models

Models channel →