Using JEV as judge: LLM citations never fully check out, ~8 sources per task

pizzababa21 · reddit · 2026-09-22

The author used JEV as a judge to check whether LLM-cited sources actually support the claims in web search retrieval tasks, with a strength score. Models typically cite around 8 sources per task, and no test run had all 8 pass. With plain search and no rules, JEV even caught outdated information by noticing article dates months old despite the user asking for recent updates. The most common failure it caught: statements drawing on multiple sources while citing only one. The author is optimistic this will make next-gen models considerably more reliable via better-vetted AI-generated test data.

Original post →

More from Models

Models channel →