Mira-Scene solves generative 3D scene layout with pixel-aligned coordinate maps, +39.8% 3D-IoU

Yang-Tian Sun · hf · 2026-09-22

Placing generated 3D objects into coherent scene layouts remains hard: holistic methods sacrifice object detail, while compositional methods rely on sparse, unbounded pose variables that generalize poorly. Mira-Scene is a compositional 3D scene reconstruction framework that replaces sparse pose regression with dense, bounded correspondence recovery.

Its core is the Canonical Coordinate Map (CCM), a pixel-aligned field mapping every visible object pixel to a surface coordinate in the object's bounded canonical space. Paired with a monocular scene-space Point Cloud Map, CCM yields dense canonical-to-scene correspondences from which transforms are recovered via robust geometric alignment—trainable from scalable object-level 3D data without scene-level layout annotations. A multimodal diffusion transformer with modality-specific expert streams jointly generates geometry and CCMs.

Across indoor, outdoor, synthetic, and in-the-wild scenes, Mira-Scene outperforms strong baselines with relative gains of 39.8% in 3D-IoU and 16.5% in 2D-IoU over SAM3D, using limited open-source training data.

Original post →

More from Research

Research channel →