VGGT-Diff routes visual geometry into video diffusion for sparse-view novel view synthesis

XPENG-AI · hf · 2026-09-29

VGGT-Diff routes VGGT-Ω geometry latents into a pretrained video diffusion model via a confidence-aware router, with Point-Track Residual Consistency for multi-view stability, achieving competitive or SOTA results on sparse-view NVS interpolation and extrapolation.

Original post →

More from Multimodal

Multimodal channel →