Critique of LoCoMo and LongMemEval: Seeking better memory benchmarks
LowDistribution3995 · reddit · 2026-08-25
The author criticizes current AI memory benchmarks like LoCoMo and LongMemEval for significant flaws: LoCoMo has poor answer keys and subjective constraints leading to irreproducible high scores, while LongMemEval is just a repetitive needle-in-a-haystack test. Simple FAISS or even Ctrl+F can outperform them. The author seeks alternative benchmarks that measure recall accuracy in a meaningful way, as existing ones use unrealistic third-party conversation data.
More from Research
- ARC Prize 2026: Tufa Labs open-sources winning solution — fchollet · 2026-08-25
- Extremely fast sine approximation with -120dB harmonics — teropa · 2026-08-25
- Pew study of 500K pages: over a third published since ChatGPT show AI-written text — The Decoder · 2026-08-25
- Microsoft Research releases Skala 1.1 for improved computational chemistry accuracy — vdbergrianne · 2026-08-25
- Research Shows Strong RL Prompts Don't Need Detail; Conciseness Helps — xeophon · 2026-08-25
- Case study: Two LLMs missed a future-data bug in coding and review loop — niacolhealth · 2026-08-25