PixARMesh rebuilds a full editable 3D scene from one photo

TheZachMueller · x · 2026-07-20

A CVPR 2026 paper from UC San Diego and Lambda, **PixARMesh**, turns a single photo into a fully editable 3D scene. - The scene is represented as a single token sequence and generated autoregressively. - It avoids SDFs, occupancy fields, and multi-stage layout optimization. - Reported scene-level F-Score on 3D-FRONT: **33.55%** vs **25.00%** for DepR. - Chamfer Distance improves from **0.153** to **0.099**. - Each object is produced as an artist-ready mesh with a few thousand faces. The method uses pixel-aligned image features plus global scene context, then predicts object poses and meshes token by token in one forward pass. The post frames it as useful for robotics, self-driving scene reconstruction, game dev, and interior design.

Original post →

More from Multimodal

Multimodal channel →