HKU team's SceneMosaic generates simulation-ready scene variants 24x faster with near-zero collisions

机器之心 · wechat · 2026-09-28

Scene generation for robot simulation faces a tradeoff: agentic text-to-3D is slow (SAGE takes 7.56 hours per scene) while parametric image-to-3D is fast but physically unreliable (20-26% collision rates). A team from the University of Hong Kong proposes SceneMosaic, open-sourced, addressing speed, physics, and diversity.

Method:

Results: On SceneEval-100, semantic quality matches the best baseline (POS 84.3 vs 83.5), 24x faster than SceneSmith, and a variant scene costs only 0.03 hours. Relation retention is 99.1% vs 71.5% for Gaussian perturbation at equal diversity.

Limitation: relies on a single input image, so currently room-scale only.

Original post →

More from Embodied

Embodied channel →