PixARMesh rebuilds a full editable 3D scene from one photo
TheZachMueller · x · 2026-07-20
A CVPR 2026 paper from UC San Diego and Lambda, PixARMesh, turns a single photo into a fully editable 3D scene.
- The scene is represented as a single token sequence and generated autoregressively.
- It avoids SDFs, occupancy fields, and multi-stage layout optimization.
- Reported scene-level F-Score on 3D-FRONT: 33.55% vs 25.00% for DepR.
- Chamfer Distance improves from 0.153 to 0.099.
- Each object is produced as an artist-ready mesh with a few thousand faces.
The method uses pixel-aligned image features plus global scene context, then predicts object poses and meshes token by token in one forward pass. The post frames it as useful for robotics, self-driving scene reconstruction, game dev, and interior design.
More from Multimodal
- Midjourney style code share: --sref 2912175708 — tisch_eins · 2026-09-11
- The full prompt-to-3D-game workflow: Hyper3D Rodin MCP plus Codex, no reference image — FellMentKE · 2026-09-11
- Building a 3D landing page with GPT-6 Astra and Hyper3D Rodin MCP, no modeling needed — FellMentKE · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- MiniMax H3 MAX nails cooking anime clips: 15-second curry demo with prompts shared — Hailuo_AI · 2026-09-11
- MiniMax Music Production Toolkit 2.5 for ComfyUI adds full mastering chain — Vivid_Promise1700 · 2026-09-11