Remember-R1: Tackling 'Visual Forgetting' in Multimodal Long-Chain Reasoning

新智元 · wechat · 2026-08-08

Multimodal models often suffer from 'distributional visual forgetting' during long-chain reasoning: as generated text grows, the model's reliance on the image degrades, shifting to language priors. To address this, researchers introduced Remember-R1.

Core Approach: Instead of altering the inference process, it integrates three process rewards during RL post-training to force the model to continuously invoke visual evidence throughout the reasoning chain:

Results: Remember-R1-7B improves across seven benchmarks. For instance, MathVista increased by 7.5 points and MMVet by 12.93 points, without sacrificing base visual perception. This research offers significant post-training optimization insights for long-horizon agents and long-video understanding tasks.

Original post →

More from Research

Research channel →