NVIDIA's UNREAL paper lets one LLM both retrieve and answer, lifting recall from 49% to 73%
mark_k · x · 2026-10-08
Most AI systems use a separate retrieval model to find documents, then hand them to an LLM. NVIDIA's UNREAL paper shows the LLM itself can do both: it uses its own internal representations to select relevant information from an index, adding fewer than 500,000 trainable parameters while keeping the base model frozen.
- On a 3-billion-token Wikipedia index, complete-evidence retrieval recall rises from 49.1% to 73.2% on HotpotQA, and from 31.7% to 60.1% on 2WikiMultiHopQA.
- The same mechanism filters distractors from long prompts: NoLiMa accuracy at 128K tokens jumps from 1% to 24.83%.
- It also cuts compute and time-to-first-token from roughly 32K tokens onward.
One model finding the evidence and answering could simplify knowledge-heavy AI systems considerably.
Related event: NVIDIA's UNREAL Unifies Retrieval and Long-Context in a Single Model(3 posts)→
More from Infra
- DRAM contract prices quadruple in 9 months, but makers realize gains very differently — tengyanAI · 2026-10-08
- Abusers hop across inference providers, so providers must coordinate evictions — natolambert · 2026-10-08
- Enterprise AI trends toward model routers, not one giant model — ingliguori · 2026-10-08
- Crowdsourced harness x model benchmark: 3090 beats 5090 with qwen3.8-flash-next config — dh7net · 2026-10-08
- llama.cpp distributes inference across heterogeneous devices: MiMo 2.6 Flash at 40 tok/s over 10 GbE — joao_gante · 2026-10-08
- Java-based jitLLM claims 90% of llama.cpp perf on NVIDIA GPUs via TornadoVM CUDA compilation — mikebmx1 · 2026-10-08