MM-R2 adds agentic planning to multimodal RAG before retrieval starts
_reachsumit · x · 2026-07-28
This paper proposes MM-R2, an agentic multimodal RAG framework that reasons about what to retrieve and where to search before retrieval.
- The authors argue that many multimodal RAG systems retrieve directly from raw image-text inputs over a flat evidence space.
- That design struggles when the user intent is under-specified and the evidence space is unstructured.
- MM-R2 first builds an intent-grounded retrieval state from the image-question pair, capturing the information need, grounded referent, and retrieval constraints.
- It then searches a structured KnowledgeMap, selecting relevant retrieval units before issuing grounded queries.
- To support training, the team created MM-R2-Traj, a large-scale dataset of multi-step retrieval trajectories.
- They train the system with supervised fine-tuning and GRPO.
- On Infoseek and Encyclopedic VQA, the paper reports better accuracy and more interpretable retrieval traces.
More from Research
- NASA/JPL to Host 2026 OPERA Workshop on GeoAI — giswqs · 2026-08-27
- Inverse-designed quantum emitters enable ultralow-power nonlinear activations — bravo_abad · 2026-08-27
- Hugging Face investigation highlights lack of oversight for AI swarms — deanwball · 2026-08-27
- CMU team uses physics-based models to bridge neuroscience and robotics — lukas_m_ziegler · 2026-08-27
- GoogleDeepMind, Partners Validate PySyft Privacy Computing in Practice — iamtrask · 2026-08-27
- Tampering One Web Page Skews AI Recommendations 27% of the Time — YvesMulkers · 2026-08-27