DeepMind's Scaffolding Minds Gains +5.6 Points Across Nine Visual Reasoning Benchmarks

GoogleDeepMind · hf · 2026-09-30

Scaffolding Minds: Optimizing Latent Representations for Multimodal Reasoning

Background: Latent reasoning enables visual chain-of-thought via a two-stage paradigm: SFT with latent tokens encoding a helper image, then RL refinement with reward feedback.

Two limitations identified:

Method: Learn a dedicated scaffolding encoder providing an optimized latent-space target, and learn both the mean and variance of the RL sampler. The two improvements are complementary.

Results: +9.5 points over the strongest latent-reasoning baseline on FrozenLake spatial planning, widening to +19 on 32x32 grids, and +5.6 points on average across nine visual reasoning benchmarks.

Original post →

More from Research

Research channel →