VoxMem benchmark: none of 15 audio LLMs top 40% on spoken multi-session memory
unimelb-hf · hf · 2026-09-30
University of Melbourne introduces VoxMem, a benchmark for multimodal memory in large audio language models: 3,196 instances over 34,743 spoken sessions (177 hours), spanning four acoustic evidence types (semantics, speaker identity, paralinguistic cues, environmental sound) and four memory operations (extraction, multi-session reasoning, temporal tracking, refusal), across 8K-64K token context budgets.
- Evaluating 15 LALMs, no model exceeds 40% at 32K context.
- Models remember what was said far better than who said it, how, or what was audible—info recoverable only from audio.
- The gap widens with harder operations and longer histories, with qualitatively distinct failure modes per evidence type.
- VoxMem provides a taxonomy (acoustic evidence × memory operations) as a foundation for measuring spoken conversational memory progress.
More from Research
- ImageJevBench v0.2 launches on 573 images, Imajev 4B tops five-model ranking — airesearch12 · 2026-09-30
- alphaXiv ships OpenResearch desktop app to turn coding agents into research agents — gekobraa · 2026-09-30
- Government weather model WRF ported to GPUs, running an order of magnitude faster with 250m fog forecasts — Scobleizer · 2026-09-30
- François Fleuret: AI math will dwarf human math, focus on lean proofs not explainability — francoisfleuret · 2026-09-30
- SortedRL: Microsoft Research tackles 70-74% GPU idle time in LLM reinforcement learning — burkov · 2026-09-30
- UMass professor Luc Rey-Bellet's stochastic processes lecture notes on Markov chains and MCMC — michaelchchoi · 2026-09-30