Do Audio LLMs Listen Before They Act? Diagnosing Acoustic-Context Gating in Voice Agents
The authors introduce VGBench, a 1,018-item diagnostic benchmark for action-level addressedness across side-talk, self-talk, and speaker-switch scenarios. Using this benchmark, they find that raw Audio LLMs and training-free adaptations often identify the target tool yet rarely withhold action under speaker shifts, with the highest raw switch mute rate being 14%. They use VoxGate as a post-training case study, demonstrating that supervised training mutes 91.3% of switched commands while choosing the correct tool for nearby wearer commands. An exploratory GRPO stage further improves side-talk accuracy and self-talk muting.
VGBench is a 1,018-item diagnostic benchmark measuring multi-cue acoustic-context gating.
Raw Audio LLMs achieve a highest raw switch mute rate of 14%.
Supervised training in VoxGate mutes 91.3% of switched commands while choosing the correct tool for nearby wearer commands.