Do Audio LLMs Listen Before They Act? New Benchmark Shows Raw Models Mute Only 14% of Bystander Commands
Yanjie Zhang · hf · 2026-10-02
VGBench is a 1,018-item diagnostic benchmark for action-level addressedness in voice agents across side-talk, self-talk, and speaker-switch scenarios, with a shared action space of silence, tool call, and natural-language answer. Six raw Audio LLMs and three training-free adaptations often identify the target tool but rarely withhold action on bystander commands—the highest raw switch mute rate is just 14%. A post-training case study, VoxGate, mutes 91.3% of switched commands while choosing correct tools for nearby wearer commands; an exploratory GRPO stage raises side-talk accuracy from 68.4% to 70.9% and self-talk muting from 52.0% to 60.0%. Factorized controls show the benchmark measures multi-cue acoustic-context gating rather than isolated speaker identity.
More from Research
- AMap open-sources ABot-Recon: streaming 3D reconstruction from video with a 12-frame local context — rsasaki0109 · 2026-10-02
- Neuralink pretrains decoders on 50,000+ hours of neural data—dataset may be the real moat — CurieuxExplorer · 2026-10-02
- CyberGym cybersecurity benchmark effectively saturated on verified task subset — aryaman2020 · 2026-10-02
- SemEval-2027 Task 9 calls for teams on 4-language multimodal news framing analysis — preslav_nakov · 2026-10-02
- Alternating prompt and model upgrades lift science agent from 42% to 73% — rohanpaul_ai · 2026-10-02
- ISMIR paper teaches a transformer to play in 12 jazz piano legends' styles — umpedronosapato · 2026-10-02