Jev as an LLM judge flops: scores nearly everything positively, disagrees with humans
amplifiedamp · x · 2026-09-18
A developer tried using Jev as a scorer in OntBench, a benchmark for how well different LLMs build ontology maps, and found it doesn't work: it scores almost everything very positively and disagrees with both the author's manual ratings and Codex's ratings (all scorers blinded).\n\nThe author shared the full scoring prompt — including evidence boundary instructions and a five-level rubric — asking the community whether something is wrong with the setup.
More from Models
- Self-Proclaimed ChatGPT Co-Inventor Launches Jev, Claims 200x Speed at 1/400 Cost — iamrobotbear · 2026-09-18
- Codex Pro User Says Usage Limits Got 5-10x Worse, Can't Even Buy Another Plan — Junra · 2026-09-18
- GPT-6 Astra Deciphers an Undeciphered 1918 German WWI Radio Transmission — moultano · 2026-09-18
- Sakana AI Introduces Fugu Max and Fugu Ultra v2 Models — SakanaAILabs · 2026-09-18
- Gemini 3.8 Live Architecture Breakdown: Sub-100ms Native Audio and Real-Time Tool Calling — 4bTechDecode · 2026-09-18
- Anthropic Opens Life Sciences Verification Program, Unlocks Mythos for Biologists — EricBuess · 2026-09-18