Uni-LaDiR unifies image, text and 3D reasoning via latent diffusion thoughts
Lianhuiq · x · 2026-10-06
Researchers introduce Uni-LaDiR (Unified Latent Diffusion Reasoner), a unified reasoning framework for both VLMs and VLAs.
The core idea: humans don't reason separately in images, text, and 3D point clouds — thinking should live in a more abstract latent space regardless of whether the CoT data comes from which modality. Uni-LaDiR uses diffusion to generate latent thoughts in that shared space, making reasoning independent of modality-specific representations.
More from Research
- 11-Square Packing Optimality Proved and Formalized in Lean With Help From Astra and Claude — ctjlewis · 2026-10-07
- Researcher unveils RSI paradigm: generic disposable agents plus an evolving knowledge base — yisongyue · 2026-10-07
- John Urschel proves Gaussian elimination growth factor is ~n^1/2, settling a Trefethen conjecture — ctjlewis · 2026-10-07
- Where Does Memory Live? RNNs, Transformers, and SSMs Compared Through Working Memory — Pretty_Upstairs9035 · 2026-10-07
- KNOWS Benchmark Jointly Evaluates Agents on Search, Tools and Visual Understanding — anmarasovic · 2026-10-07
- ID-Forcing Keeps Long Video Generation In-Distribution, Enabling Minute-Scale Videos Without Fine-Tuning — _akhaliq · 2026-10-07