CVPR 2025 paper learns 3D spatial audio from unlabeled video using camera ego-motion
量子位 · wechat · 2026-07-26
Researchers from Tsinghua and the University of Michigan presented a CVPR 2025 Highlight paper that learns 3D spatial audio perception from unlabeled in-the-wild video.
Core idea
- Instead of expensive 360° microphones or manual labels, the method uses camera ego-motion as a free supervision signal.
- The model watches paired audio/video clips, estimates camera rotation and translation from the visual frames, and learns sound-source spatial shifts from those motion labels.
Data and setup
- The team built two datasets: iPhoneFountain and iPhoneMusicDataset.
- Training uses adjacent audio snippets plus corresponding video frames.
- Ego-motion labels are encoded into simple left/right and front/back style targets for supervision.
Results and takeaways
- The framework performs strongly on both the new datasets and the public Fountain benchmark.
- Ablations suggest distance to the sound source is a major source of bias: farther sounds become harder to localize.
- The paper’s broader claim is that real-world internet video can serve as scalable training data for spatial audio, with potential applications in robotics, AR/VR, and multimodal learning.
More from Embodied
- AI glasses are getting spatial workspaces with 3DOF tracking — Scobleizer · 2026-07-26
- REK Tokyo says Japan’s first humanoid robot show is a success — cixliv · 2026-07-26
- MiniCPM-Robot gets attention for contextual memory and fully local tracking — aliscodes · 2026-07-26
- Musk’s Optimus gets a meme-worthy “hotel cleaning mode” demo — 4KTV · 2026-07-26
- Instagram bans Meta glasses videos used for prank harassment and covert filming — Polymarket · 2026-07-26
- Sunday Robotics says its beta robot is ready for real homes after six months — mon__lim · 2026-07-26