Do Audio LLMs Listen Before They Act? New Benchmark Shows Raw Models Mute Only 14% of Bystander Commands

Yanjie Zhang · hf · 2026-10-02

VGBench is a 1,018-item diagnostic benchmark for action-level addressedness in voice agents across side-talk, self-talk, and speaker-switch scenarios, with a shared action space of silence, tool call, and natural-language answer. Six raw Audio LLMs and three training-free adaptations often identify the target tool but rarely withhold action on bystander commands—the highest raw switch mute rate is just 14%. A post-training case study, VoxGate, mutes 91.3% of switched commands while choosing correct tools for nearby wearer commands; an exploratory GRPO stage raises side-talk accuracy from 68.4% to 70.9% and self-talk muting from 52.0% to 60.0%. Factorized controls show the benchmark measures multi-cue acoustic-context gating rather than isolated speaker identity.

Original post →

More from Research

Research channel →