OmniEcho: Spatial Audio Benchmark and Model for Embodied Agents
PKU-VaLuE-Lab · hf · 2026-09-25
PKU-VaLuE-Lab introduces OmniEchoBench and the OmniEcho model for spatial audio understanding in embodied agents.
- Benchmark: 6 tasks over 197 real-world spatial audio-visual scenes, 2,972 QA pairs, and 900 navigation samples with first-order ambisonics (FOA) audio from 30 real environments
- Pipeline: A controllable spatial audio rendering pipeline preserves geometric consistency among sound sources, visuals, and agent trajectories
- Model: OmniEcho adds an FOA spatial encoder to a pretrained semantic audio pathway, achieving SOTA on spatial audio-visual perception; sound-guided navigation approaches traditional vision-language navigation
- Open challenges: fine-grained spatial localization and distance estimation remain difficult
More from Embodied
- HRI 2027 Adds Archival Industry White Paper Track for Real-World Robot Deployments — petitegeek · 2026-09-25
- Google dev imagines orchestrating AI agents via smart glasses: tap their shoulder, watch their screen — jason_mayes · 2026-09-25
- Climbing analysis in 3D with iPhone LiDAR: open-source SAM 3.1 + ViTPose demos — dosco · 2026-09-25
- Menlo open-sources Asimov 1 humanoid locomotion policy and Isaac Lab training code — freelerobot · 2026-09-25
- Human-Robot Dialogue workshop at IROS 2026: NVIDIA, MIT, Georgia Tech speakers lined up — dhadfieldmenell · 2026-09-25
- AI glasses shipments up 263% in H1; Zhiyuan delivers 20,000th robot — 创业邦 · 2026-09-25