UCSD's WM-VLM adds an internal world model to VLMs, boosting spatial reasoning by up to 39.25 points

UCSanDiego · hf · 2026-10-06

WM-VLM: Giving VLMs an Internal World Model for Spatial Reasoning

Humans solve spatial problems by mentally simulating visual transformations, while conventional VLMs mostly reason in language. Researchers at UCSD propose WM-VLM, which attaches a lightweight world model branch to a pretrained VLM to generate intermediate visual states, enabling reasoning over both text and generated images.

Method

Results

The authors argue internal world models are a promising path toward VLMs that reason in both language and visual space.

Related event: UCSD's WM-VLM Boosts VLM Spatial Reasoning via Internal World Model(2 posts)→

Original post →

More from Multimodal

Multimodal channel →