NVIDIA's UNREAL paper lets one LLM both retrieve and answer, lifting recall from 49% to 73%

mark_k · x · 2026-10-08

Most AI systems use a separate retrieval model to find documents, then hand them to an LLM. NVIDIA's UNREAL paper shows the LLM itself can do both: it uses its own internal representations to select relevant information from an index, adding fewer than 500,000 trainable parameters while keeping the base model frozen.

One model finding the evidence and answering could simplify knowledge-heavy AI systems considerably.

Related event: NVIDIA's UNREAL Unifies Retrieval and Long-Context in a Single Model(3 posts)→

Original post →

More from Infra

Infra channel →