Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering
Jaewoo Jung, Hyeonseo Yu, Honggyu An, Jisang Han, Mungyeom Kim, Minkyeong Jeon, Heeseong Shin, Wonjun Moon, Federico Tombari, Daniel Barath, Marc Pollefeys, Seungryong Kim, Sunghwan Hong
NeurIPS 2026
cs.CV, cs.CL
2026-09-30
Learnable Gaussian summary tokens let a 7B MLLM reconstruct multi-view scenes as compact 3D Gaussians before answering, hitting 68.5 on SPAR-Bench, 29 points above a 72B model.
Multimodal LLMs (MLLMs) are strong on a single image and weak on a pile of views of the same scene. Ask a question that requires knowing where the camera sits relative to the furniture, and even frontier models fall well short of human performance. Two fixes dominate the literature: stitching 3D point coordinates or bird's-eye-view markers into patch features, or fusing in features from reconstruction foundation models like VGGT and CUT3R. Both feed the model pixel-level geometry, and both have been delivering shrinking returns.
The KAIST AI and ETH Zürich team starts from a different observation. Cognitive science says humans do not keep pixel-accurate correspondences in their heads. Given several views, people recognize the same objects across them, infer rough relative camera geometry, and assemble a coarse 3D layout. The representation behind human spatial reasoning is object-level and approximate. So the question becomes: instead of telling the model the geometry, can it be made to imagine the scene?
Imagine3D-LLM builds on LLaVA-Video-7B. The change sits in the input sequence: 2,592 learnable Gaussian summary tokens inserted between the image tokens and the text tokens. With 32 input images producing 6,720 image tokens, that is roughly 81 summary tokens per image. Training combines three objectives:
Two design choices carry the paper. First, Gaussians are predicted from dedicated summary tokens rather than directly from image features, which blocks the shortcut of copy-pasting 2D content into image-aligned Gaussians. Second, the summary tokens are far fewer than the image tokens (2,592 versus 6,720), an information bottleneck: an object appearing in many views has to collapse into a shared token, so the model must work out which objects recur and how they fit together in 3D. K-means over the trained summary tokens clusters them by object, with no clustering supervision anywhere.
The ablation nails down the causality: feeding the teacher's query tokens straight into the LLM scores 56.4 on SQA3D EM, distillation without reconstruction 56.6, both sitting at the baseline's 56.5. Only the reconstruction loss moves the number, to 63.8.
| Benchmark | Metric | Imagine3D-LLM-7B | Best prior |
| SQA3D | EM | 63.8 | Ross3D 63.0 |
| ScanQA | CIDEr | 109.3 | Ross3D 107.0 |
| Scan2Cap | [email protected] | 99.2 | 3DRS 86.1 |
| ScanRefer | [email protected] | 62.8 | 58.3 (same-recipe baseline) |
| SPAR-Bench | average | 68.5 | 3DThinker-7B 63.3 |
SPAR-Bench splits into low, medium, and high cognitive levels, and the gains concentrate in medium (view-change inference) and high (spatial imagination). A controlled baseline sharing the same backbone, data, and schedule goes from 51.5 to 67.0 on the medium split (+15.5), so the gain traces to the objective rather than the data. On average, the 7B model beats Qwen2.5-VL-72B by 29.1 points.
ScanQA EM is one metric it does not win (29.9 against Ross3D's 30.8). Dropping distillation and training reconstruction-only for 4 epochs reaches 63.7/67.9 versus the full model's 63.8/68.5 at 1 epoch, so distillation mostly buys training time. The token count has a sweet spot: 1,296 is too few, 5,184 loosens the bottleneck, and both underperform 2,592.
Table 7 is the actionable finding for anyone building 3D MLLMs: handing the LLM features from a 3D foundation model does essentially nothing, while making the LLM compute the 3D structure itself does. That is a direct counterpoint to the passive-geometry line of positional embeddings, visual markers, and feature fusion. And although reconstruction supervision touches only the summary tokens, the LLM's own image features come out more view-consistent, with sharper cross-view attention and object-consistent PCA structure. A structured objective reshapes representations across the stack. Inference cost is modest: 2,592 extra tokens, no teacher, no point cloud input. For general MLLM teams this is a portable training objective, not a bespoke architecture.