CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — evaluation 12 upvotes

Do Audio LLMs Listen Before They Act? Diagnosing Acoustic-Context Gating in Voice Agents

QUESTION — How do current Audio LLMs handle acoustic-context gating for action execution in voice agents?

The authors introduce VGBench, a 1,018-item diagnostic benchmark for action-level addressedness across side-talk, self-talk, and speaker-switch scenarios. Using this benchmark, they find that raw Audio LLMs and training-free adaptations often identify the target tool yet rarely withhold action under speaker shifts, with the highest raw switch mute rate being 14%. They use VoxGate as a post-training case study, demonstrating that supervised training mutes 91.3% of switched commands while choosing the correct tool for nearby wearer commands. An exploratory GRPO stage further improves side-talk accuracy and self-talk muting.

VGBench is a 1,018-item diagnostic benchmark measuring multi-cue acoustic-context gating.

Raw Audio LLMs achieve a highest raw switch mute rate of 14%.

Supervised training in VoxGate mutes 91.3% of switched commands while choosing the correct tool for nearby wearer commands.

Yushi98 · 26 Sept 2026 read the original ↗
↑