CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — evaluation 126 upvotes

VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models

QUESTION — How to rigorously benchmark multimodal memory in large audio language models across diverse acoustic evidence types and multi-session conversational histories?

The authors argue that existing spoken conversational benchmarks ignore non-lexical audio evidence and multi-session histories. They propose a taxonomy combining acoustic evidence types with memory operations, and introduce VoxMem: 3,196 evaluation instances across 34,743 spoken sessions (177 hours) spanning context budgets from 8K to 64K tokens. Evaluating 15 Large Audio Language Models (LALMs), they find that no model exceeds 40% accuracy at 32K tokens, and models struggle significantly with speaker identity, paralinguistic cues, and environmental sounds compared to lexical content.

VoxMem comprises 3,196 evaluation instances over 34,743 spoken sessions totaling 177 hours.

The benchmark crosses four acoustic evidence types with four memory operations across context budgets from 8K to 64K tokens.

Evaluating 15 LALMs, no model exceeds 40% accuracy at 32K tokens.

AustinXiao · 26 Sept 2026 read the original ↗
↑