VRRL: RL with visual feedback teaches VLMs to self-correct their mistakes

gregd_nlp · x · 2026-09-04

Researchers introduce VRRL, an RL method addressing the fact that vision-language models fail to correct mistakes even when the visual environment shows what went wrong. VRRL trains VLMs to self-reflect on visual feedback, achieving stronger OOD accuracy in visually grounded reasoning than RL and SFT baselines.

Original post →

More from Research

Research channel →