Critique of LoCoMo and LongMemEval: Seeking better memory benchmarks

LowDistribution3995 · reddit · 2026-08-25

The author criticizes current AI memory benchmarks like LoCoMo and LongMemEval for significant flaws: LoCoMo has poor answer keys and subjective constraints leading to irreproducible high scores, while LongMemEval is just a repetitive needle-in-a-haystack test. Simple FAISS or even Ctrl+F can outperform them. The author seeks alternative benchmarks that measure recall accuracy in a meaningful way, as existing ones use unrealistic third-party conversation data.

Original post →

More from Research

Research channel →