DistScene: single-image compositional 3D scene generation with explicit environment modeling
Kunming Luo · hf · 2026-10-07
DistScene generates compositional 3D scenes from a single image. Unlike methods treating scenes as object collections, it models the environment as an explicit scene component providing geometric context for object placement. Three components:
- Scene-Frame Generation: jointly generates environment and object components in a shared coordinate frame so geometry and placement are learned together;
- Object-Centric Refinement: refines each object in a local frame with scene context;
- Object-to-Scene Distillation: transfers pretrained object-generation priors to scene generation via automatically composed and rendered synthetic scenes.
Indoor and outdoor benchmarks show improved scene-level spatial coherence over baselines.
More from Multimodal
- How Claude Opus 5.5 'Makes Videos': It Writes Programs That Draw, Not Frames — dotey · 2026-10-07
- ComfyUI Qwen img2img enhancer adds per-region reference mask control — Capitan01R- · 2026-10-07
- DRAMA 1.0 launches: edit only the expression layer, swap 10 emotions in existing footage — nikola_mr64990 · 2026-10-07
- video-shotcraft hits 10.5k stars: cinematic product videos via Claude Code + Remotion — tom_doerr · 2026-10-07
- Third-party test: Google's Nano Banana 2.1 beats ChatGPT Images, 5x faster and cheaper — prajdabre · 2026-10-07
- ByteDance's DuoMatching: few-step video generation wins 80%+ human preference — ByteDance · 2026-10-07