NVIDIA's Axolotl3D does occlusion-aware 3D shape completion from multimodal inputs
NVIDIAAI · x · 2026-09-17
NVIDIA's Spatial Intelligence Lab introduced Axolotl3D at ECCV 2026, a multimodal, occlusion-aware 3D shape completion model.
Approach
- Unlike prior 3D generative models assuming single-view, fully visible inputs, Axolotl3D jointly conditions on images, visibility masks, camera parameters, and a partial point cloud
- The point cloud acts as a geometric anchor for faithful completion; camera parameters keep multi-view outputs aligned in a shared 3D coordinate system
- A unified training strategy synthesizes diverse conditioning regimes from large-scale 3D data, fine-tuning Hunyuan3D-DiT and decoding via ShapeVAE
Results
- State-of-the-art on Toys4K and OmniObject3D in both clean and occluded settings
- Strong real-world reconstruction and geometry-consistent editing results
More from Multimodal
- Quiver AI launches Arrow 2, faster and higher-quality editable vector graphics generation — stuffyokodraws · 2026-09-17
- QuiverAI launches image generation models Arrow 2 and Arrow 2 Telos — stuffyokodraws · 2026-09-17
- QuiverAI launches Arrow 2, its most advanced model for editable vector graphics — stuffyokodraws · 2026-09-17
- SuperSplat now shows splat counts for your 3D scenes — willeastcott · 2026-09-17
- ai-toolkit, the open-source LoRA trainer, lands on Pinokio for 1-click installs — cocktailpeanut · 2026-09-17
- Spotify's AI Persona Badge Targets Photorealistic Identities, Not Royalties — mixtapedmonk · 2026-09-17