Single Image to Full 3D Scene: Adaptive Chunking Extends Object Generators to Outdoor Rome
Jiraphon Yenphraphai · hf · 2026-10-07
Researchers repurpose object-centric 3D generators (e.g., Trellis 2) to generate complete 3D scene meshes—including unobserved surfaces—from a single image, for both indoor and outdoor scenes.
Key techniques:
- Adaptive chunking: scene partitions scale with camera distance—small chunks nearby preserve detail, large chunks cover distant structures like buildings;
- Explicit 2D-3D correspondence: lifted image features with awareness of free space, observed surfaces, and unobserved regions;
- 4,000 synthesized outdoor scenes broaden training beyond indoor-heavy datasets.
The method outperforms all baselines in geometric accuracy and perceptual quality on Tanks and Temples, ScanNet++, and in-the-wild images.
More from Multimodal
- DRAMA 1.0 launches: edit only the expression layer, swap 10 emotions in existing footage — nikola_mr64990 · 2026-10-07
- video-shotcraft hits 10.5k stars: cinematic product videos via Claude Code + Remotion — tom_doerr · 2026-10-07
- Third-party test: Google's Nano Banana 2.1 beats ChatGPT Images, 5x faster and cheaper — prajdabre · 2026-10-07
- Video World Models Flunk Physics: Best Model Scores 57.76/100 on New 40-Task Benchmark — Mingju Gao · 2026-10-07
- ByteDance's DuoMatching: few-step video generation wins 80%+ human preference — ByteDance · 2026-10-07
- DistScene: single-image compositional 3D scene generation with explicit environment modeling — Kunming Luo · 2026-10-07