Agent memory API benchmarks are broken: self-reported vs third-party scores differ by 20-30 points

Efficient_Joke3384 · reddit · 2026-09-04

A developer comparing agent memory APIs (Mem0, Zep, Letta, MemoryLake) found that everyone cites the LoCoMo benchmark, but self-reported and third-party numbers diverge wildly:

Scores for the same product on the same benchmark swing 20-30 points depending on who's reporting. The vendors don't trust each other either — Zep published a post directly questioning Mem0's numbers.

The author concludes that differing test setups and correctness criteria have turned "X% on LoCoMo" into a marketing line rather than a comparable metric, and asks whether any benchmark in this space is actually trustworthy.

Original post →

More from coding & agent

coding & agent channel →