VoxMem benchmark: none of 15 audio LLMs top 40% on spoken multi-session memory

unimelb-hf · hf · 2026-09-30

University of Melbourne introduces VoxMem, a benchmark for multimodal memory in large audio language models: 3,196 instances over 34,743 spoken sessions (177 hours), spanning four acoustic evidence types (semantics, speaker identity, paralinguistic cues, environmental sound) and four memory operations (extraction, multi-session reasoning, temporal tracking, refusal), across 8K-64K token context budgets.

Original post →

More from Research

Research channel →