VRRL: RL with visual feedback teaches VLMs to self-correct their mistakes
gregd_nlp · x · 2026-09-04
Researchers introduce VRRL, an RL method addressing the fact that vision-language models fail to correct mistakes even when the visual environment shows what went wrong. VRRL trains VLMs to self-reflect on visual feedback, achieving stronger OOD accuracy in visually grounded reasoning than RL and SFT baselines.
More from Research
- Google uses AI to map full fruit fly brain, reconstructing 166,000+ neurons in 3D — burny_tech · 2026-09-04
- Language models can control their own attention: 52% decoding cost cut on Gemma 4 31B — jm_alexia · 2026-09-04
- Pedro Domingos proposes scrapping accept/reject: reviewers rate papers, attendees vote on what gets in — pmddomingos · 2026-09-04
- Classic Reward Hacking: A Soccer Bot That Started Running Backwards Loops for No Reason — generativist · 2026-09-04
- Baseten Launches Base Labs, an Open Research Lab for Open-Source Frontier AI — baseten · 2026-09-04
- New protocol enables private, verifiable LLM inference outsourcing with 179MB local storage — chaumian · 2026-09-04