Devils in Question Relay: Source-Conditioned Relay Steering to Mitigate Hallucinations in Audio-visual Large Language Models
QUESTION — How can we mitigate source-confused grounding hallucinations caused by cross-modal interference in audio-visual large language models?
The authors identify a question-relay mechanism in audio-visual large language models where question states carry interfering cues from unused modalities. To fix this, they propose SECRET, a training-free method that steers question representations toward required-source evidence using pathway interventions. Experiments on CMM and AVHBench benchmarks across three AVLLMs show that SECRET mitigates source-confused grounding hallucinations by up to +18.0 and +7.1 percentage points over base models.
SECRET consistently outperforms prior training-free methods, substantially mitigating source-confused grounding hallucinations (e.g., up to +18.0 and +7.1 percentage points over base models).
271754echo · 29 Sept 2026
read the original ↗