Imagine3D-LLM teaches MLLMs to imagine a 3D scene before answering

kaist-ai · hf · 2026-10-01

KAIST's Imagine3D-LLM takes inspiration from human spatial reasoning: identify common objects across views, infer relative viewpoint geometry, and assemble a coarse 3D layout.

The model appends learnable summary tokens after image tokens, decodes them into a compact 3D Gaussian Splatting representation supervised by photometric reconstruction, and trains jointly with next-token prediction. Notably, reconstruction supervision on the summary tokens alone induces stronger cross-frame correspondence in the LLM's image features, suggesting 3D-aware signals propagate model-wide. Imagine3D-LLM consistently beats prior approaches on spatial reasoning and 3D understanding benchmarks.

Original post →

More from Multimodal

Multimodal channel →