Study: broad question coverage beats repeated reads in agentic RAG evaluation budgets
_reachsumit · x · 2026-10-06
An arXiv paper studies how to allocate evaluation token budgets in agentic RAG across questions, search trajectories, and repeated reads.
- Experiments on HotpotQA and MuSiQue measure allocation precision, reading efficiency, and cost boundaries via retrieval-feedback comparisons.
- At matched cost (34.1–34.4M model tokens), broader question coverage lowers standard error by 33% versus five reads and 12.6% versus three trajectories.
- Archived nested and question-only forecasts predict optimal allocations within 4.0% and 3.5%.
- One-read variance penalties vs the fitted optimum are 0–9.9%; at search prices of $0–1 per 1,000 requests, more questions beat more trajectories.
- Temperature zero cuts answer disagreement from 14.3% to 3.4% with similar comparison precision.
Related event: Study: More Questions Beats More Reads in Agentic RAG Eval Budgets(2 posts)→
More from Research
- Temperature 0 doesn't make LLMs deterministic: 1,000 runs of Qwen3-235B yielded 80 outputs — lmoroney · 2026-10-08
- macro2mind Trains LLMs on Prediction Markets to Simulate Individuals, +15.5 Points Zero-Shot — youjiaxuan · 2026-10-08
- FractAL introduces soft acquisition-strategy selection for batch-mode active learning — anshulkundaje · 2026-10-08
- DeLM's decentralized multi-agent system runs 2.49x faster, but MAS evals pick wildly different metrics — jyangballin · 2026-10-08
- Study: RAG retrieval diversification helps only on redundant multi-evidence pools, paper proposes per-query rule — _reachsumit · 2026-10-08
- Quantize by Drift: label-free mixed-precision quantization for text embedders hits 0.911 Spearman — _reachsumit · 2026-10-08