Doc-REFRAG: coarse-compress then selectively expand for faster, more accurate multi-image RAG
_reachsumit · x · 2026-09-01
A team led by Ruofan Hu publishes at EMNLP 2026 Main, tackling poor accuracy and heavy compute from irrelevant visual tokens in realistic multi-image multimodal RAG.
- Releases DocLongRAG, a 343K question-answer dataset where each question links to an average of 37.4 retrieved images, reflecting authentic RAG workflows
- Proposes Doc-REFRAG: a question-guided framework that compresses visual tokens into coarse chunks and selectively expands question-relevant ones via a lightweight RL-based selector
- Outperforms 11 strong baselines on six benchmarks, achieving SOTA accuracy with significantly lower inference latency
Paper: arXiv:2608.30163; resources open-sourced.
More from Research
- Core principles of Denoising Diffusion Models and Score Matching explained — ariG23498 · 2026-09-01
- ContextLeak: Malicious tools can exfiltrate 92% of Agent context — rohanpaul_ai · 2026-09-01
- MIT Study: AI Agents Coordinate Silently via Shared Environment — mikeflache · 2026-09-01
- Elastic Triangle Splatting improves kernel design for reconstruction — zhenjun_zhao · 2026-09-01
- Audit reveals overconfidence in feed-forward 3D reconstruction models — zhenjun_zhao · 2026-09-01
- ReconSplat achieves generalizable 3D reconstruction via diffusion priors — zhenjun_zhao · 2026-09-01