NVIDIA's UNREAL: one model unifies corpus retrieval and long-context at 128K+
nvidia · hf · 2026-10-07
- Idea: use a frozen LLM's internal representations for both corpus-level retrieval and long-context evidence selection, adding <500K trainable parameters with the backbone untouched.
- Retrieval: on a 3B-token, 21M-chunk Wikipedia index, UNREAL beats SOTA retriever-reranker systems — HotpotQA recall jumps from 49.1% to 73.2%, 2WikiMultiHopQA from 31.7% to 60.1%, MuSiQue from 8.8% to 14.4%.
- Long context: the same mechanism prunes distractors before generation, lifting NoLiMa accuracy from 1.0% to 24.83% at 128K tokens and LV-Eval F1 from 49.97% to 54.66% at 256K.
- Efficiency: from 32K tokens onward it beats full-context inference on FLOPs and time-to-first-token, with gains growing with context length.
Related event: NVIDIA Unveils UNREAL: One Model Unifies Retrieval and Long Context(2 posts)→
More from Infra
- One of the Last American Chestnut Groves to Be Destroyed for a Data Center — Promptmethus · 2026-10-07
- CtrlCache Speeds Up Interactive Video World Models 1.21–1.41x Without Retraining — Shangye Song · 2026-10-07
- The 2019 Mac Pro with 1.5TB RAM would be the ultimate local LLM machine today — Odd-Capital-847 · 2026-10-07
- Team claims sub-5-second full weight sync for 1T-parameter RL training — saurabh_shah2 · 2026-10-07
- SlimWise prunes MoE experts only at decode, boosting throughput up to 1.81x — Gunho Park · 2026-10-07
- Ora scanned 107,797 sites: average agent-readiness score just 41/100 — EdenEmarco177 · 2026-10-07