VGGT-Diff routes visual geometry into video diffusion for sparse-view novel view synthesis
XPENG-AI · hf · 2026-09-29
VGGT-Diff routes VGGT-Ω geometry latents into a pretrained video diffusion model via a confidence-aware router, with Point-Track Residual Consistency for multi-view stability, achieving competitive or SOTA results on sparse-view NVS interpolation and extrapolation.
More from Multimodal
- One prompt, no cuts: Kling 4.0 video realism is getting hard to distinguish from real footage — FinanceYF5 · 2026-09-29
- One prompt, no cuts: Kling 4.0 video realism is getting hard to distinguish from real footage — FinanceYF5 · 2026-09-29
- A cyberpunk city built entirely in code: thousands of towers, live-synthesized soundtrack, zero audio files — techartist_ · 2026-09-29
- Atlases Are Already Inside: Recovering Population Templates by Making Diffusion Models Collapse — kwangmoo_yi · 2026-09-29
- EvolvingAvatar Uses Test-Time Training to Make 3D Talking Heads Adapt as Conversations Unfold — HFUT-AI · 2026-09-29
- SciGen-Verifier brings explainable, reasoning-driven verification to scientific image generation — Jiali Chen · 2026-09-29