DeepMind's Scaffolding Minds Gains +5.6 Points Across Nine Visual Reasoning Benchmarks
GoogleDeepMind · hf · 2026-09-30
Scaffolding Minds: Optimizing Latent Representations for Multimodal Reasoning
Background: Latent reasoning enables visual chain-of-thought via a two-stage paradigm: SFT with latent tokens encoding a helper image, then RL refinement with reward feedback.
Two limitations identified:
- The SFT stage uses an off-the-shelf vision encoder, yielding latent representations poorly aligned with downstream reasoning
- Existing RL only applies deterministic regularization, constraining drift without enabling exploration of alternative latent trajectories
Method: Learn a dedicated scaffolding encoder providing an optimized latent-space target, and learn both the mean and variance of the RL sampler. The two improvements are complementary.
Results: +9.5 points over the strongest latent-reasoning baseline on FrozenLake spatial planning, widening to +19 on 32x32 grids, and +5.6 points on average across nine visual reasoning benchmarks.
More from Research
- Diffusion Models Tutorial Accepted to NeurIPS 2026 Alongside 7 Paper Acceptances — mittu1204 · 2026-09-30
- AMB3R-SLAM: Kilometer-Scale Real-Time SLAM on One Consumer GPU, Cutting ATE by 70% — rsasaki0109 · 2026-09-30
- Prefix-Reuse FLOPs: new metric exposes hidden cost of arbitrary context edits in LLM serving — RulinShao · 2026-09-30
- Tencent Hunyuan releases ExplorationBench to measure how AI systems explore — TencentHunyuan · 2026-09-30
- Fully open MolmoAct 2 tops independent robotics benchmark LIBERO-MAX on dynamic robustness — DJiafei · 2026-09-30
- CompVis improves Distributional Diffusion Models: 4.48 FID at 4 steps on ImageNet — CompVis · 2026-09-30