VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models
The authors argue that existing spoken conversational benchmarks ignore non-lexical audio evidence and multi-session histories. They propose a taxonomy combining acoustic evidence types with memory operations, and introduce VoxMem: 3,196 evaluation instances across 34,743 spoken sessions (177 hours) spanning context budgets from 8K to 64K tokens. Evaluating 15 Large Audio Language Models (LALMs), they find that no model exceeds 40% accuracy at 32K tokens, and models struggle significantly with speaker identity, paralinguistic cues, and environmental sounds compared to lexical content.
VoxMem comprises 3,196 evaluation instances over 34,743 spoken sessions totaling 177 hours.
The benchmark crosses four acoustic evidence types with four memory operations across context budgets from 8K to 64K tokens.
Evaluating 15 LALMs, no model exceeds 40% accuracy at 32K tokens.