VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis
Kangjie Chen, Xiangyu Li, Dongbin Zhang, Chaoda Zheng, Shijia Chen, Jinhao Deng, Hongbin Lin, Choo Sin Wai, Minqi Wang, Minghao Yang, Dake Zhong, Guorui Song, Yu Zhang, Xianming Liu, Boyang Wang
cs.CV, cs.AI
2026-09-27
Routes VGGT-Omega 3D points and confidence into a Wan2.1 video diffusion prior; DL3DV PSNR 18.104, +0.92 dB over same-backbone FrameCrafter, using only ~1K training scenes.
Sparse-view novel view synthesis (NVS) has long been split in two. Reconstruction methods (NeRF, Gaussian splatting, feed-forward models like LVSM) keep observed geometry faithful, but once the query camera moves away from the sources, unobserved regions are under-constrained and predictions degenerate into patch-like smears biased toward source colors. Diffusion methods bring strong generative priors to complete unseen content, but the source-to-query correspondence stays implicit, and the prior often overrides geometry-supported evidence: structural drift, unstable occlusions, cross-view inconsistency.
This paper (XPeng Motors, CUHK, Tsinghua) welds the two together: a visual geometry foundation model, VGGT-Ω, supplies explicit 3D points, correspondences, and confidences to ground a pretrained video diffusion model, Wan2.1-I2V-14B.
The operative idea is geometry as a condition, not as a scene reconstruction. Three components:
Two engineering choices round it out. During training the routed geometry condition is randomly dropped (10%) or attenuated (20%), so the model never treats imperfect projections as ground truth. The same dropout enables a Geometry-Prior CFG at inference that strengthens query-geometry guidance during early denoising steps.
On DL3DV, 6,188 targets, six source views, trained on only 980 scenes (AnySplat in the same table used 254K):
| Method | PSNR | SSIM | LPIPS | DreamSim |
| LVSM | 17.090 | 0.478 | 0.333 | 0.204 |
| SEVA | 16.150 | 0.470 | 0.253 | 0.088 |
| FrameCrafter (same Wan2.1-14B) | 17.180 | 0.445 | 0.223 | 0.066 |
| VGGT-Diff | 18.104 | 0.508 | 0.222 | 0.062 |
The overall SSIM lead belongs to the regression model DepthSplat at 0.566. The cleanest comparison is FrameCrafter, same backbone and comparable data: +0.924 dB PSNR, +0.063 SSIM. Zero-shot on Mip-NeRF 360, VGGT-Diff ranks first in LPIPS and DreamSim and second in PSNR (E-RayZer leads, with far worse perceptual errors).
Stratified by pose difficulty (interpolation/extrapolation × near/mid/far): first in five of six bins, trailing SEVA by 0.020 dB in interpolation-near, with margins over FrameCrafter between 0.582 and 1.228 dB. Cross-view consistency is probed by reconstructing eight jointly generated views with two frozen reconstructors: Chamfer-L1 down 26.7% (VGGT-Ω) and 16.4% (Pi3, never used in training or conditioning) versus FrameCrafter. Two independent judges agreeing rules out shared-backbone bias.
In the ablations, removing all VGGT-Ω-dependent components costs 1.730 dB, the largest single effect; removing PTRC costs 0.438 dB; condition regularization 0.302 dB; inference-time CFG adds only 0.115 dB. Swapping PTRC for direct velocity matching drops PSNR from 18.581 to 17.769, so the gain comes from the residual formulation itself. Two appendix results: scaling to 10K scenes and 100K steps raises PSNR from 18.79 to 20.28 with no plateau yet, and running with only 3 or 4 source views (trained with 6, no retraining) still beats FrameCrafter on all four metrics.
For anyone building 3D generation or reconstruction pipelines, this is direct evidence that geometry foundation model features, with 3D points and confidence attached, can serve as explicit diffusion conditions without first reconstructing a complete 3D scene. The data efficiency is real: about 1K scenes plus a borrowed Wan prior, reproducible for a small team. PTRC's align-the-residual-instead-of-the-prediction form transfers to other multi-view generation tasks.
The authors' own list: cross-view consistency is measured against point clouds reconstructed from GT views rather than physical scans, so it is a relative measure; joint-denoising memory limits the view count and calls for chunked generation; only static scenes are handled; jointly fine-tuning VGGT-Ω actually loses 0.647 dB, so the geometry encoder only works frozen.
Reading closely, three more: inference cost of a 14B backbone with 50 sampling steps is never discussed; the aggregate protocol follows FrameCrafter, and regression methods still top PSNR and SSIM on Mip-NeRF 360, so cross-camp conclusions should lean on the stratified numbers; the headline SOTA claim holds only for LPIPS and DreamSim on Mip-NeRF 360. The paper is a preprint and has not been peer-reviewed.