LoG-VGGT uses cross-window attention for memory-efficient long-sequence 3D reconstruction
zhenjun_zhao · x · 2026-09-22
LoG-VGGT balances local temporal modeling with global camera consistency: cross-window attention at a subset of transformer layers keeps memory bounded, while camera tokens cross-attend to compact register tokens enforce scene-level constraints, improving depth accuracy and long-horizon pose stability across multiple benchmarks.
More from Embodied
- Snap Specs' $2,195 Launch Faces 'Creep Glasses' Backlash and Apple's $1,999 Foldable iPhone — adariostrange · 2026-09-22
- Huawei MateBook Fold review: a 1.16kg 13-inch notebook that unfolds into an 18-inch 3.3K OLED, ~$3,700, China-only — anthara_ai · 2026-09-22
- Gemini beats GPT-6 Astra at robot capture the flag, winning 70% of matches — chris_j_paxton · 2026-09-22
- Meta Ray-Ban glasses talking to Muse could become the best consumer AI app overnight — ChrisUniverse · 2026-09-22
- JHU to host Scalable Tactile Sensing for Dexterous Manipulation workshop at IROS 2026 — _krishna_murthy · 2026-09-22
- SLIM-init: line-feature VIO initialization for degenerate motions, accepted to IROS 2026 — zhenjun_zhao · 2026-09-22