RoMa-Ω: Swapping RoMa v2's DINOv3 Backbone for VGGT-Ω Sets New Image Matching SOTA
kwangmoo_yi · x · 2026-09-11
Researchers from Chalmers and collaborators released RoMa-Ω (ECCVW 2026), probing what feed-forward 3D models know about image matching.
- The change is minimal: replace RoMa v2's DINOv3 backbone with VGGT-Ω features and retrain.
- The result outperforms RoMa and RoMa v2 across many benchmarks, achieving state-of-the-art matching — notably beating RoMa on the difficult WxBS and HardMatch benchmarks where RoMa v2 fell short.
- The model remains a slow but powerful dense matcher; code and paper are available on GitHub.
Related event: RoMa-Ω swaps DINOv3 for feed-forward 3D features, sets matching records(3 posts)→
More from Research
- DeepSeek's CED vs GLM's KV reuse: a developer unpacks how the cache-sharing modes actually differ — stochasticchasm · 2026-09-11
- SG-JEPA Paper: World Models That Train on Earth and Deploy on Mars, Halving Zero-Shot Physics Error — randall_balestr · 2026-09-11
- Code-as-Policy article explores general models learning to operate robots like software — yawnxyz · 2026-09-11
- ICML Generative AI and Creativity workshop releases new survey paper — lasha_nlp · 2026-09-11
- AVSplat: assist-view preconditioning fixes dense-view degradation in feed-forward 3DGS — zhenjun_zhao · 2026-09-11
- RRSI descriptor enables radiation, rotation and scale-invariant multimodal image matching — zhenjun_zhao · 2026-09-11