Amazon's ReaLVR finds latent visual tokens ignore image evidence, scales latent reasoning to 235B
amazon · hf · 2026-10-01
Latent visual reasoning (LVR) lets multimodal LLMs compute in continuous latent tokens instead of verbalizing every step—but latent tokens aren't directly observable, making them hard to supervise.
- Finding: the authors identify a latent evidence-credit gap—latent tokens respond only weakly to image perturbations that change the correct answer, traced to lack of explicit supervision in GRPO training.
- Method: ReaLVR brings visual-evidence supervision to free-running latent trajectories, contrasting correct vs model-generated wrong answers and relevant vs mismatched evidence.
- Results: it outperforms LVR baselines across three model families (63.7% five-task average on Qwen2.5-VL-7B) and is the first to scale latent visual reasoning to frontier scale (235B) with robust gains.
More from Research
- Hugging Face open-sources Tau, a readable terminal coding agent built to teach — mervenoyann · 2026-10-01
- Researcher: new models solve old problems, but learning theory lacks predictive conjectures — brianryhuang · 2026-10-01
- ParallelPilot paper: 63% higher throughput for parallel coding agents — erichorvitz · 2026-10-01
- Multi-harness RL guide: LFM2.5 jumps 42% to 54% with 31% fewer tool calls — SergioPaniego · 2026-10-01
- LATENT wins IROS 2026 award: humanoid robots rally at human level from imperfect motion data — chris_j_paxton · 2026-10-01
- Bocconi paper: teach causal reasoning in the age of LLMs — daveholtz · 2026-10-01