VGGT-Prime: Compute-Adaptive Mixture-of-Heads for Efficient Visual Geometry Transformers
Abteen Arab, Guile Wu, Chengjie Huang, Dongfeng Bai
cs.CV
2026-09-21
Huawei Noah routes VGGT's 384 global-attention heads to mean pooling, surrogate attention, or full softmax by saliency, claiming 8× speed on 1000-frame sequences and up to 14× with token merging.
Feed-forward visual geometry models such as VGGT reconstruct cameras, depth, and point maps from multi-view images in one pass. Global attention over all tokens from all views is what makes that work, and it is also why cost grows quadratically with the number of frames. Long sequences get slow fast.
Most accelerators attack token redundancy: FastVGGT merges tokens, SparseVGGT sparsifies keys and values, HTTM merges tokens per head. They never ask whether VGGT's 24 global-attention layers, 16 heads each, 384 heads in total, are all doing useful geometric work.
Huawei Noah scores each head with maximum integrated gradients (MIG) and maximum post-softmax attention (MaxAttn). The two rankings agree at Spearman ρ=0.741. Pruning heads from least to most salient barely hurts at first, then collapses once the high-saliency heads go. Heads cluster into high (ranks about 1–100), medium (101–250), and low (251–384).
Attention maps match that split. Low-saliency heads have high entropy and low query diversity; they behave like mean pooling. Medium heads stay somewhat selective but look similar across queries. High-saliency heads are low-entropy and query-specific; they carry cross-view correspondence and need full softmax.
VGGT-Prime puts a lightweight router on each head. It attention-pools the head's queries into one vector, attends that vector to all keys, and takes the maximum as saliency sh. Thresholds τlow=0.01 and τhigh=0.20 are fit on a CO3Dv2 validation split and then frozen.
Training is two-stage distillation. The router and surrogate first match a frozen VGGT teacher's latents, then the surrogate and full-attention branches are fine-tuned with soft routing. Mean pooling is attached last, with no extra training.
On ScanNet-500, VGGT takes 90.1 s at Chamfer 0.442; Prime takes 19.2 s at 0.441, about 4.7×. On dense 7-Scenes (stride 3) both hit CD 0.115, while time drops from 38.1 s to 9.8 s (3.9×). On ETH3D, Prime's CD 0.99 beats VGGT's 1.07.
| Setting | VGGT | Prime |
| ScanNet-500 CD / s | 0.442 / 90.1 | 0.441 / 19.2 |
| Dense 7-Scenes CD / s | 0.115 / 38.1 | 0.115 / 9.8 |
| 7-Scenes AUC@30° | 79.8 | 78.6 |
| CO3Dv2 RTA | 97.3 | 99.2 |
| Sintel AbsRel | 0.592 | 0.524 |
The abstract's 8× figure is for 1000-frame sequences against an already accelerated VGGT variant from MapAnything. Combined with FastVGGT token merging, speedup is about 11× at 75% merge and nearly 14× at 90%. The same routing applies to VGGT-Ω, π³, and Depth Anything 3, at roughly 2× on sparse 7-Scenes.
Ablations are harsh. Dropping the residual correction cuts CO3Dv2 AUC@30 from 87.2 to 42.6. Binary keep-or-drop routing blows 7-Scenes CD from 0.115 to 0.581. Replacing the surrogate with linear attention is both slower and worse.
This is an acceleration axis orthogonal to token reduction. Teams already running FastVGGT can stack the two. If the job is feed-forward reconstruction on hundreds to thousands of frames, this is more useful than another sparse-attention variant. The router is cheap; the savings come from skipping full softmax on most heads.
It is an engineering improvement, not a new geometric representation. Quality stays close to VGGT, with small wins on ETH3D CD, Sintel depth, and CO3Dv2 RTA, and a small drop in pose AUC.
The 8× number is against Keetha et al.'s accelerated VGGT, not a naive original. HTTM and HeSS numbers come from the authors' PyTorch reimplementations. Thresholds are locked on CO3Dv2; the appendix says routing distributions transfer, but the search itself is not adaptive. Distillation needs a frozen teacher and staged training. Mean-pooling the weakest heads may hide damage on highly repetitive texture, which the main tables do not isolate.