CVPR 2025 paper learns 3D spatial audio from unlabeled video using camera ego-motion

量子位 · wechat · 2026-07-26

Researchers from Tsinghua and the University of Michigan presented a CVPR 2025 Highlight paper that learns 3D spatial audio perception from unlabeled in-the-wild video.

Core idea

Data and setup

Results and takeaways

Original post →

More from Embodied

Embodied channel →