Agent memory API benchmarks are broken: self-reported vs third-party scores differ by 20-30 points
Efficient_Joke3384 · reddit · 2026-09-04
A developer comparing agent memory APIs (Mem0, Zep, Letta, MemoryLake) found that everyone cites the LoCoMo benchmark, but self-reported and third-party numbers diverge wildly:
- Mem0: 92.5% self-reported vs 64% in third-party tables
- Zep: 94.7% self-reported vs 85% third-party
- Letta: no self-reported number, 74% third-party
- MemoryLake: claims 94% with no independent verification
Scores for the same product on the same benchmark swing 20-30 points depending on who's reporting. The vendors don't trust each other either — Zep published a post directly questioning Mem0's numbers.
The author concludes that differing test setups and correctness criteria have turned "X% on LoCoMo" into a marketing line rather than a comparable metric, and asks whether any benchmark in this space is actually trustworthy.
More from coding & agent
- Running Qwen3.8-Flash-Next with 256K context at 16 tok/s on DDR4 and a Tesla T4 — BusTiny207 · 2026-09-04
- Handy Omarchy plugin wraps Claude CLI with a menu bar UI and improved interface — DanWahlin · 2026-09-04
- Scheduled Agent Stopped After One Weak Run: A Lesson in Defining Search Depth — daani_maas · 2026-09-04
- xAI Launches Early-Beta Grok Bot: AI Teammates That Log Into Your Tools and Finish Work — aitrendz_xyz · 2026-09-04
- Claude burns 24h building cells that take seconds to measure, in overnight benchmark fail — StefanoGogioso · 2026-09-04
- Connect Claude Code to NotebookLM via MCP to read full docs without burning tokens — udmrzn · 2026-09-04