Doc-REFRAG: coarse-compress then selectively expand for faster, more accurate multi-image RAG

_reachsumit · x · 2026-09-01

A team led by Ruofan Hu publishes at EMNLP 2026 Main, tackling poor accuracy and heavy compute from irrelevant visual tokens in realistic multi-image multimodal RAG.

Paper: arXiv:2608.30163; resources open-sourced.

Original post →

More from Research

Research channel →