Jev fails as an LLM scorer on OntBench: rates almost everything positively, contradicting human and Codex ratings
amplifiedamp · x · 2026-09-18
A developer tried using Jev as an automated scorer in OntBench (a benchmark for how well different LLMs build ontology maps) and found it unusable: it scores almost everything very positively, disagreeing with both the author's manual ratings and Codex's ratings.
The full five-level scoring prompt (from 'multiple major mistakes' to 'supported claims with explicit limits') is shared, asking the community what went wrong. It's a textbook LLM-as-judge failure: the scorer lacks discrimination and skews positive.
More from Models
- Self-Proclaimed ChatGPT Co-Inventor Launches Jev, Claims 200x Speed at 1/400 Cost — iamrobotbear · 2026-09-18
- Codex Pro User Says Usage Limits Got 5-10x Worse, Can't Even Buy Another Plan — Junra · 2026-09-18
- GPT-6 Astra Deciphers an Undeciphered 1918 German WWI Radio Transmission — moultano · 2026-09-18
- Sakana AI Introduces Fugu Max and Fugu Ultra v2 Models — SakanaAILabs · 2026-09-18
- Gemini 3.8 Live Architecture Breakdown: Sub-100ms Native Audio and Real-Time Tool Calling — 4bTechDecode · 2026-09-18
- Anthropic Opens Life Sciences Verification Program, Unlocks Mythos for Biologists — EricBuess · 2026-09-18