VGGT-Prime: compute-adaptive mixture-of-heads slashes redundancy in visual geometry transformers

zhenjun_zhao · x · 2026-09-22

VGGT-Prime accelerates feed-forward visual geometry models like VGGT by targeting architectural redundancy rather than token redundancy, showing only a subset of global-attention heads carries critical geometric information.

A lightweight router estimates the appropriate compute level per head and dynamically assigns it to mean pooling, surrogate attention, or full softmax attention, cutting the quadratic cost of growing view counts while keeping competitive reconstruction quality across multiple datasets.

Original post →

More from Research

Research channel →