ST-AudioLM From Sony AI and POSTECH Tracks Moving Sound Sources With 40 Trajectory Tokens
mittu1204 · x · 2026-10-01
Researchers from POSTECH and Sony AI released ST-AudioLM (EMNLP 2026 Main), with code and demos now public.
- Current audio-language models treat clips as global events; ST-AudioLM instead binds semantics to motion using one semantic token plus 40 trajectory tokens, unifying understanding, spatial localization, and trajectory modeling
- It ships with ST-AudioQA, a benchmark built from first-order ambisonic (FOA) renderings of static and moving sources, with metadata for identity, direction, distance, and motion enabling multi-source grounding and compositional QA
- A time-resolved FOA encoder (ST-Audio Encoder) learns event semantics alongside source trajectories
- Binaural demos include an approaching ambulance siren and a receding hair dryer; the work was led by Sony AI intern Hyun-Bin Oh
More from Multimodal
- ThinkV2V: Reasoning-Driven Video Editing — a 5B Model Beats 10B Baselines on Complex Instructions — Donghao Zhou · 2026-10-01
- Fudan's IDSpect Decomposes Chinese Characters into Radicals for Fine-Grained Text-to-Image RL Rewards — Fudan-University · 2026-10-01
- Imagine3D-LLM teaches MLLMs to imagine a 3D scene before answering — kaist-ai · 2026-10-01
- JHU unveils PowerSim: differentiable physics that simulates and re-renders captured 3D scenes — anand_bhattad · 2026-10-01
- 7000 Midjourney Images + Suno Track + Opus 5.5: 2-Hour AI Creation Cost Just $6.80 — ciguleva · 2026-10-01
- Free Dot avatar maker: 25 shapes, 12 materials, 32 styles, no sign-up — yihui_indie · 2026-10-01