Remember-R1: Tackling 'Visual Forgetting' in Multimodal Long-Chain Reasoning
新智元 · wechat · 2026-08-08
Multimodal models often suffer from 'distributional visual forgetting' during long-chain reasoning: as generated text grows, the model's reliance on the image degrades, shifting to language priors. To address this, researchers introduced Remember-R1.
Core Approach: Instead of altering the inference process, it integrates three process rewards during RL post-training to force the model to continuously invoke visual evidence throughout the reasoning chain:
- Visual Vocabulary Reward: Constrains the model to consistently mention key visual phrases in subsequent reasoning.
- Visual Memory Reward: Penalizes rapid drops in visual attention during later reasoning stages by analyzing attention mechanisms.
- Critical Region Reward: Ensures the model's attention remains focused on image regions genuinely relevant to the question.
Results: Remember-R1-7B improves across seven benchmarks. For instance, MathVista increased by 7.5 points and MMVet by 12.93 points, without sacrificing base visual perception. This research offers significant post-training optimization insights for long-horizon agents and long-video understanding tasks.
More from Research
- Prove2Me: Claude agents wrote 13M lines of Lean in 11 days to formalize Fermat's Last Theorem — liuzhuang1234 · 2026-09-21
- EvoOntology: a self-evolving ontology layer bridges the agent-data gap — RUC-DataLab · 2026-09-21
- Xiaomi's CodeMidas builds 5,545 coding RL environments from raw source code — XiaomiMiMo · 2026-09-21
- GraphSkillEvo evolves graph-structured skills for LLM agents, +4% on benchmarks — Rui Sun · 2026-09-21
- DeformSmith generates physics-grounded deformable assets for robot manipulation — Can Li · 2026-09-21
- RewardAI launches OM-1 robot foundation model trained on human data alone, zero-shot across arms and humanoids — zipengfu · 2026-09-21