GenIA: Rendering-Guided Test-Time Alignment Turns SAM3D into SOTA Image-to-3D

JonathonLuiten · x · 2026-10-09

Researchers from Tübingen AI Center and Meta Reality Labs introduce GenIA, fixing the misalignment problem of image-to-3D generative models. Existing methods are either not pixel-aligned (wrong colors, missing details) or generalize poorly. GenIA grounds a frozen SAM3D foundation model at test time using differentiable rendering guidance during denoising — no retraining needed. It improves object pose via geometry-derived translation/scale, aligns appearance through visibility-biased attention and cross-observation fusion, and supports multi-view and monocular video inputs, recovering canonical appearance for dynamic objects. It achieves state-of-the-art results on both seen and unseen object parts across synthetic and real benchmarks.

Original post →

More from Multimodal

Multimodal channel →