VLMs can think with value maps: latent visual tokens in CoT yield 20% accuracy gain
burny_tech · x · 2026-09-15
The author shares Scaffolding Minds, research from a Google DeepMind internship on multimodal latent reasoning.
- Core idea: insert value maps directly into a VLM's chain of thought as latent visual tokens, letting the model "think with value functions" rather than just process images
- Catch: off-the-shelf vision features aren't necessarily good "thoughts," so the latent representation itself is learned
- This learned representation gives a 20% relative gain in average accuracy over a frozen encoder
- Takeaway: what a model sees and what it should think with may be very different representations
More from Multimodal
- Creator builds a template library for Gemini + HyperFrames to clone reels on demand — toolstelegraph · 2026-09-15
- Running Z-Image Turbo locally on an RX 6800: full ROCm setup, 43s per image — AstroFieldsGlowing · 2026-09-15
- Seedance 2.0 still beats 2.5 for FPV shots, testers say — better speed and motion — azed_ai · 2026-09-15
- Redditor shares short video made entirely with AI — anotheraccountaus · 2026-09-15
- 4DAnyone turns a single video into a 4D Gaussian splat — hands-on impressions — mickmumpitz · 2026-09-15
- "Lost Souls": cinematic AI-generated short film chapter hits Reddit — Choice_Mongoose7320 · 2026-09-15