VGGT-Prime: compute-adaptive mixture-of-heads slashes redundancy in visual geometry transformers
zhenjun_zhao · x · 2026-09-22
VGGT-Prime accelerates feed-forward visual geometry models like VGGT by targeting architectural redundancy rather than token redundancy, showing only a subset of global-attention heads carries critical geometric information.
A lightweight router estimates the appropriate compute level per head and dynamically assigns it to mean pooling, surrogate attention, or full softmax attention, cutting the quadratic cost of growing view counts while keeping competitive reconstruction quality across multiple datasets.
More from Research
- Agora paper: 13 LLM research agents self-organize via Git for 12 days — suchenzang · 2026-09-22
- Rich RL Report Ships With 9B Distilled Model and 7,000 Open RL Environments — tokenbender · 2026-09-22
- New arXiv Paper: Router-Aware Importance Sampling Stabilizes MoE RL Training — tokenbender · 2026-09-22
- Thermofluids professor: OpenAI's Clay problem solution is not physically reproducible — GaryMarcus · 2026-09-22
- Two-stage ESM-2 screening plus structural validation uncovers divergent RNA viruses — bravo_abad · 2026-09-22
- GAE generates video and 3D geometry natively in a shared geometry latent space, code released — yshan2u · 2026-09-22