Using an LLM as benchmark scorer fails: over-optimistic ratings diverge from human judgment
amplifiedamp · x · 2026-09-18
The author tried using Jev as a scorer on OntBench, a benchmark for how well different LLMs build ontology maps — and it just doesn't work. It scores almost everything very positively, disagreeing with both the author's manual ratings and Codex's blinded ratings.
The author shared their scoring prompt asking whether they misused it, noting that Jev now tracks Codex fairly closely overall but remains overly optimistic about everything.
More from Models
- Self-Proclaimed ChatGPT Co-Inventor Launches Jev, Claims 200x Speed at 1/400 Cost — iamrobotbear · 2026-09-18
- Codex Pro User Says Usage Limits Got 5-10x Worse, Can't Even Buy Another Plan — Junra · 2026-09-18
- GPT-6 Astra Deciphers an Undeciphered 1918 German WWI Radio Transmission — moultano · 2026-09-18
- Sakana AI Introduces Fugu Max and Fugu Ultra v2 Models — SakanaAILabs · 2026-09-18
- Gemini 3.8 Live Architecture Breakdown: Sub-100ms Native Audio and Real-Time Tool Calling — 4bTechDecode · 2026-09-18
- Anthropic Opens Life Sciences Verification Program, Unlocks Mythos for Biologists — EricBuess · 2026-09-18