Agentic RAG eval budgets: broader question coverage cuts standard error 33% vs repeated reads

CarnegieMellonU · hf · 2026-10-08

This study measures optimal evaluation budget allocation across questions, search trajectories, and repeated reads in agentic RAG using HotpotQA and MuSiQue. At 34M model tokens, broader question coverage lowers standard error by 33% versus five reads and 12.6% versus three trajectories. Archived forecasts predict allocations within 4%. Under recorded fees, more questions beat more trajectories at search prices of $0–1 per 1,000 requests; temperature zero cuts answer disagreement from 14.3% to 3.4% with similar comparison precision.

Related event: Study: More Questions Beats More Reads in Agentic RAG Eval Budgets(2 posts)→

Original post →

More from coding & agent

coding & agent channel →