RoMa-Ω: Swapping DINOv3 for VGGT-Ω Features Reveals What Feed-Forward 3D Models Know About Matching
zhenjun_zhao · x · 2026-09-11
New paper RoMa-Ω asks what feed-forward 3D reconstruction models know about image matching.
- Analyzes three scenarios: zero-shot patch-feature matching, direct matching of 3D point predictions, and training a full matcher on top of the learned representations
- Findings: feed-forward models like VGGT perform poorly at zero-shot matching (especially in later layers), yet provide strong representations for linear probing and full matching pipelines; even untrained raw predictions enable competitive matching under moderate viewpoint changes and modality gaps
- Building on this, the authors retrain RoMa v2 by replacing its DINO backbone with VGGT-Ω features
Related event: RoMa-Ω swaps DINOv3 for feed-forward 3D features, sets matching records(3 posts)→
More from Research
- Geoffrey Irving fails to prove Lean kernel correct after weeks of attempts — geoffreyirving · 2026-09-11
- Researchers prove a single 2D billiard ball can simulate a universal Turing machine — prof_g · 2026-09-11
- New math paper: generic degree-3 JC(3) counterexamples have Galois group S3 — ctjlewis · 2026-09-11
- Mathematician volunteers to mentor three AI-assisted Mathathon student teams — soumitrashukla9 · 2026-09-11
- Recurrent denoisers: adding a persistent hidden state makes diffusion an anytime solver — burkov · 2026-09-11
- S2-Attention: hardware-aware Triton kernels make sparse attention actually fast — burkov · 2026-09-11