UCSD's WM-VLM adds an internal world model to VLMs, boosting spatial reasoning by up to 39.25 points
UCSanDiego · hf · 2026-10-06
WM-VLM: Giving VLMs an Internal World Model for Spatial Reasoning
Humans solve spatial problems by mentally simulating visual transformations, while conventional VLMs mostly reason in language. Researchers at UCSD propose WM-VLM, which attaches a lightweight world model branch to a pretrained VLM to generate intermediate visual states, enabling reasoning over both text and generated images.
Method
- Two-stage training: first learn to predict the next visual state, then learn to use it for reasoning
- Verifiable benchmarks: programmatically constructed 2D/3D mental rotation tasks with verifiable intermediate visual states, measuring both generation quality and reliance on the generated states
Results
- Consistently outperforms the supervised fine-tuned backbone on 2D and 3D mental rotation, with gains up to 39.25 percentage points
- Ablations show removing or corrupting the generated visual states sharply reduces performance, confirming the gains come from the internal world model
The authors argue internal world models are a promising path toward VLMs that reason in both language and visual space.
Related event: UCSD's WM-VLM Boosts VLM Spatial Reasoning via Internal World Model(2 posts)→
More from Multimodal
- Dev uses GitHub Copilot app to generate a rebuildable 60s hype video via PR — DanWahlin · 2026-10-06
- Tencent releases new open-weight video generation model — Famous-Sport7862 · 2026-10-06
- Nano Banana 2.1 on Google Flow stuns with editorial portrait quality — full prompt shared — aziz4ai · 2026-10-06
- TagScribeR rebuilt: free local dataset studio with native LoRA training on AMD ROCm and NVIDIA — ArchAngelAries · 2026-10-06
- ComfyUI queue stuck? A maintainer's checklist to separate validation, node failures and lost progress — fluxdraw · 2026-10-06
- First try with Seedance 2.5: the model butchered the on-screen text at the end — atomantsmasher · 2026-10-06