GenIA: Rendering-Guided Test-Time Alignment Turns SAM3D into SOTA Image-to-3D
JonathonLuiten · x · 2026-10-09
Researchers from Tübingen AI Center and Meta Reality Labs introduce GenIA, fixing the misalignment problem of image-to-3D generative models. Existing methods are either not pixel-aligned (wrong colors, missing details) or generalize poorly. GenIA grounds a frozen SAM3D foundation model at test time using differentiable rendering guidance during denoising — no retraining needed. It improves object pose via geometry-derived translation/scale, aligns appearance through visibility-biased attention and cross-observation fusion, and supports multi-view and monocular video inputs, recovering canonical appearance for dynamic objects. It achieves state-of-the-art results on both seen and unseen object parts across synthetic and real benchmarks.
More from Multimodal
- This Workflow Uses Astra to Edit Video Autonomously: 20 Mins Setup, 2 Mins of Feedback per Hour — nickbaumann_ · 2026-10-09
- Kling 4.0 Flash earth-zoom effect goes viral, creator shares tweaked prompt — umesh_ai · 2026-10-09
- YarnGPT relaunches with AI speech and video dubbing for African languages — saheedniyi_02 · 2026-10-09
- Running MiniMax-H3 video gen on an 8GB VRAM laptop, now asking for text-to-image picks — Zoaloo · 2026-10-09
- Synthesia launches Syren Video: prompt-to-branded AI video with conversational editing — synthesiaIO · 2026-10-09
- Seedance 2.5 used to create a 30s AAA-style survival gameplay concept with prompt — azed_ai · 2026-10-09