Uni-LaDiR unifies multimodal reasoning in latent space with diffusion-generated latent thoughts

Lianhuiq · x · 2026-10-07

Researchers introduce Uni-LaDiR (Unified Latent Diffusion Reasoner), a framework that unifies text, images, 3D point clouds, and robot state into a single latent reasoning interface for both VLMs and VLAs. The core idea: reasoning should live in an abstract latent space independent of modality, with diffusion generating 'latent thoughts'.

Motivating example: a factory cart needs location for navigation, geometry for manipulation, and changed-image detection for inspection — complementary views of one object. Uni-LaDiR lets each task draw on shared world context, helping physical AI combine and reuse context efficiently across tasks.

Related event: Uni-LaDiR Unifies Multimodal Reasoning via Latent Diffusion(3 posts)→

Original post →

More from Embodied

Embodied channel →