Imagine3D-LLM teaches MLLMs to imagine a 3D scene before answering
kaist-ai · hf · 2026-10-01
KAIST's Imagine3D-LLM takes inspiration from human spatial reasoning: identify common objects across views, infer relative viewpoint geometry, and assemble a coarse 3D layout.
The model appends learnable summary tokens after image tokens, decodes them into a compact 3D Gaussian Splatting representation supervised by photometric reconstruction, and trains jointly with next-token prediction. Notably, reconstruction supervision on the summary tokens alone induces stronger cross-frame correspondence in the LLM's image features, suggesting 3D-aware signals propagate model-wide. Imagine3D-LLM consistently beats prior approaches on spatial reasoning and 3D understanding benchmarks.
More from Multimodal
- Seedance 2.5 demo brings anime-level dual-sword choreography into photorealistic cinema — SimplyAnnisa · 2026-10-01
- Midjourney --sref 3896456162 recreates 1970s Kodak film & disco aesthetics — michaelrabone · 2026-10-01
- A finished LoRA run doesn't mean it learned the style: a reproducible SDXL validation workflow — no3us · 2026-10-01
- Editor open-sources open-fusion-mcp: Claude builds editable motion graphics inside DaVinci Resolve — JohnnyLegion · 2026-10-01
- AI-generated clip: Raven Noire writing in her diary while listening to goth music — Street-Pound5762 · 2026-10-01
- Failed MiniMax H3 VAE detail experiment yields a useful 2X detail VAE and workflow — NoMouse9610 · 2026-10-01