China Telecom and MemTensor unveil HaluMem, first operation-level benchmark for agent memory hallucinations
jiqizhixin · x · 2026-10-09
China Telecom Research Institute and MemTensor presented HaluMem, accepted at EMNLP 2026 main conference — the first operation-level hallucination benchmark for agent memory systems.
- Agent memories fail in distinct ways: misrecording a job title while organizing a conversation, keeping stale records, or retrieving the right memory yet swapping in a wrong title when answering. Identical wrong answers can stem from entirely different failures.
- HaluMem drills evaluation into three stages — memory retrieval, memory update, and memory QA — tracing exactly where an error originates and how it propagates, rather than only checking final output.
- This decomposes 'wrong answer' into distinguishable failure modes, giving memory systems a diagnostic tool.
More from Research
- Swapping harness lifts GPT-5.6 repo migration from 6.5% to 31%, paper finds — omarsar0 · 2026-10-09
- Sakana AI's Continuous Memory Machine splits short-term compute from long-term storage — dair_ai · 2026-10-09
- Delip Rao and Chris Callison-Burch publish paper on rubrics for LLM evaluation — deliprao · 2026-10-09
- Mathematician releases AI-drafted papers combining OpenAI work with accelerated logconcave sampling research — michaelchchoi · 2026-10-09
- COLM poster: behaviors thinking models amplify aren't the ones that drive good outcomes — Jeande_d · 2026-10-09
- EngramEdit: near-perfect fact editing in LLMs via conditional memory, 3x baseline CoT accuracy — teortaxesTex · 2026-10-09