VOCA: Boosting Visual Odometry Performance Using Codec Information

rsasaki0109 · x · 2026-08-03

Camera pose estimation via Visual Odometry (VO) is critical for spatial world models. However, traditional V-SLAM systems are mostly trained on raw, uncompressed videos, while real-world hardware relies on lossy compression that introduces visual artifacts hindering tracking.

Researchers introduced VOCA, a causal stereo visual-odometry method that exploits codec information to improve tracking. It achieves state-of-the-art performance for causal VO in terms of relative trajectory error, efficiency, and absolute trajectory error.

Original post →

More from Embodied

Embodied channel →