Using JEV as judge: LLM citations never fully check out, ~8 sources per task
pizzababa21 · reddit · 2026-09-22
The author used JEV as a judge to check whether LLM-cited sources actually support the claims in web search retrieval tasks, with a strength score. Models typically cite around 8 sources per task, and no test run had all 8 pass. With plain search and no rules, JEV even caught outdated information by noticing article dates months old despite the user asking for recent updates. The most common failure it caught: statements drawing on multiple sources while citing only one. The author is optimistic this will make next-gen models considerably more reliable via better-vetted AI-generated test data.
More from Models
- Dev Praises Qwen3.8-max-preview xhigh, Ran Workloads Nearly 24 Hours on Generous Quota — jasonkneen · 2026-09-22
- Musk Announces Grok 4.7, Cited Demo Builds Interactive 3D Jet Engine in Minutes — elonmusk · 2026-09-22
- GPT-6 Sol, Luna and 'Astra Minor' reportedly spotted in Azure model config — 141_1337 · 2026-09-22
- $40M-backed Jev got an open-source API clone within three days — Nickabot · 2026-09-22
- User: hard to go back to other models after DeepSeek v4.1 Flash speed — mariofilhoml · 2026-09-22
- How Meta's Muse works, revealed by the 6.8 GB filesystem it sent me — Aeroi · 2026-09-22