HKU and VAST's Mira-Scene fixes object placement in single-image 3D scene reconstruction

jiqizhixin · x · 2026-10-08

Single-image 3D generation has mastered object-level reconstruction, but scaling to full scenes introduces a placement problem: generated objects look right yet never quite line up with the original image.

The University of Hong Kong and VAST present Mira-Scene, which introduces a pixel-aligned layout representation: it first derives a layout aligned to the source image at the pixel level, then conditions 3D generation on that layout, guaranteeing strict consistency between the reconstructed scene and the input. It can be combined with GPT-6 Astra.

Original post →

More from Multimodal

Multimodal channel →