Uni-LaDiR unifies image, text and 3D reasoning via latent diffusion thoughts

Lianhuiq · x · 2026-10-06

Researchers introduce Uni-LaDiR (Unified Latent Diffusion Reasoner), a unified reasoning framework for both VLMs and VLAs.

The core idea: humans don't reason separately in images, text, and 3D point clouds — thinking should live in a more abstract latent space regardless of whether the CoT data comes from which modality. Uni-LaDiR uses diffusion to generate latent thoughts in that shared space, making reasoning independent of modality-specific representations.

Original post →

More from Research

Research channel →