LoG-VGGT uses cross-window attention for memory-efficient long-sequence 3D reconstruction

zhenjun_zhao · x · 2026-09-22

LoG-VGGT balances local temporal modeling with global camera consistency: cross-window attention at a subset of transformer layers keeps memory bounded, while camera tokens cross-attend to compact register tokens enforce scene-level constraints, improving depth accuracy and long-horizon pose stability across multiple benchmarks.

Original post →

More from Embodied

Embodied channel →