gpt-4.1-nano as RAG judge barely beats chance: AUROC 0.603 on RAGBench
Giulio Zeloni · hf · 2026-10-07
- Context: enterprises face recurring promote/revise/block decisions on RAG systems, but metrics come from fallible LLM judges. The authors present AGO, an evidence-first quality-gate framework deployed in industrial RAG assessment.
- Design: a four-state decision model (missing data and judge errors are explicit outcomes), layered scoring (deterministic checks + local guardrails + structured LLM eval), a stratified beta-binomial gate quantifying regression risk, and mandatory meta-evaluation of the judge.
- Findings (RAGBench, N=1200 per judge): gpt-4.1-nano detects non-adherent answers at AUROC 0.603 — barely above chance — despite flawless protocol output; gpt-4o reaches 0.783 but varies 0.62–0.88 across domains.
- Gate study: under regression, the decision-grade profile cuts unsafe promotion to 22.2%–35.1% vs 29.3%–41.8% for a naive gate. Point estimates alone are not a release decision.
More from coding & agent
- English Is the New Programming Language: A Real Terraform-to-AI Handoff Story — _jaydeepkarale · 2026-10-07
- Make Claude and ChatGPT cross-review each other; humans shouldn't be the courier — sujingshen · 2026-10-07
- Open-source art animation skill: 35 styles, 9 explainer grammars, runs in Claude Code — AlchainHust · 2026-10-07
- AI Engineering Roadmap 2026: Learn Software Fundamentals First, Agents Last — techNmak · 2026-10-07
- Developer ditches Sign in with ChatGPT, says old Codex login is better in every way — lucasmeijer · 2026-10-07
- From 'software engineering is dead' to ordering 12 textbooks: a dev's 8-stage AI coding arc — mattpocockuk · 2026-10-07