Using an LLM as benchmark scorer fails: over-optimistic ratings diverge from human judgment

amplifiedamp · x · 2026-09-18

The author tried using Jev as a scorer on OntBench, a benchmark for how well different LLMs build ontology maps — and it just doesn't work. It scores almost everything very positively, disagreeing with both the author's manual ratings and Codex's blinded ratings.

The author shared their scoring prompt asking whether they misused it, noting that Jev now tracks Codex fairly closely overall but remains overly optimistic about everything.

Related event: Using Jev as an LLM judge fails: near-uniform high scores diverge from humans(3 posts)→

Original post →

More from Models

Models channel →