Mira-Scene solves generative 3D scene layout with pixel-aligned coordinate maps, +39.8% 3D-IoU
Yang-Tian Sun · hf · 2026-09-22
Placing generated 3D objects into coherent scene layouts remains hard: holistic methods sacrifice object detail, while compositional methods rely on sparse, unbounded pose variables that generalize poorly. Mira-Scene is a compositional 3D scene reconstruction framework that replaces sparse pose regression with dense, bounded correspondence recovery.
Its core is the Canonical Coordinate Map (CCM), a pixel-aligned field mapping every visible object pixel to a surface coordinate in the object's bounded canonical space. Paired with a monocular scene-space Point Cloud Map, CCM yields dense canonical-to-scene correspondences from which transforms are recovered via robust geometric alignment—trainable from scalable object-level 3D data without scene-level layout annotations. A multimodal diffusion transformer with modality-specific expert streams jointly generates geometry and CCMs.
Across indoor, outdoor, synthetic, and in-the-wild scenes, Mira-Scene outperforms strong baselines with relative gains of 39.8% in 3D-IoU and 16.5% in 2D-IoU over SAM3D, using limited open-source training data.
More from Research
- Glasshouse v0.1 launches: an open memory benchmark with 2,847 questions over a 1.97M-token conversation — True_Mongoose_7073 · 2026-09-22
- Do high-volume PIs review proportionally? Skepticism on NeurIPS reciprocal reviewing — 3scorciav · 2026-09-22
- New Decision Index benchmark runs 132,422 decisions; Jev still tops at 59.5 — victormustar · 2026-09-22
- Toby Ord estimates $20M spent on AI's millennium prize result, $200M for solid data — tobyordoxford · 2026-09-22
- Dev reverse-engineers DLSS 5 neural rendering, reimplements it bit-exact in Vulkan at 7.8ms/1080p — bdsqlsz · 2026-09-22
- The Pain Axis: research probes whether LLMs have internal pain-related patterns — voices4AI · 2026-09-22