HKU team's SceneMosaic generates simulation-ready scene variants 24x faster with near-zero collisions
机器之心 · wechat · 2026-09-28
Scene generation for robot simulation faces a tradeoff: agentic text-to-3D is slow (SAGE takes 7.56 hours per scene) while parametric image-to-3D is fast but physically unreliable (20-26% collision rates). A team from the University of Hong Kong proposes SceneMosaic, open-sourced, addressing speed, physics, and diversity.
Method:
- Reconstruct scenes with SAM3 + SAM3D into a hierarchical "scene tree" with independently evolvable local units (e.g., bookshelf + books)
- Physics-first pipeline: containment correction and gravity simulation before a Critic-Actor loop where the Critic gives qualitative advice and the Actor resolves symbolic pose expressions precisely
- Diversity via Cartesian product of local variants, then Max-Min greedy search with novelty distance
Results: On SceneEval-100, semantic quality matches the best baseline (POS 84.3 vs 83.5), 24x faster than SceneSmith, and a variant scene costs only 0.03 hours. Relation retention is 99.1% vs 71.5% for Gaussian perturbation at equal diversity.
Limitation: relies on a single input image, so currently room-scale only.
More from Embodied
- Edge AI chip startup SiMa.ai raises $150M Series C at $1.45B valuation — dauber · 2026-09-29
- Cheap DC motor with harmonic drive: outer gear rides on PTFE tubing — _Stocko_ · 2026-09-29
- C. Elegans Connectome With 302 Neurons Simulated in Real Time on a Smartwatch — bodyaz · 2026-09-29
- Cloudini 1.4 released: 20% better pointcloud compression and slightly faster — facontidavide · 2026-09-29
- LeCun Amplifies SF World Models Reading Club on Oct 10 Featuring AdaJEPA and NVIDIA Robotics — ylecun · 2026-09-29
- Embodied AI is repeating cobots' path: demos proven dead ends a decade ago — yongqianme · 2026-09-29