AI Memory Benchmarks Don't Add Up: One User Can't Tell Which Product Actually Works

Efficient_Joke3384 · reddit · 2026-09-06

A long-time user of character-chat apps and self-built memory setups says tone and emotional quality are now near-human, but memory remains the weak link — even on expensive flagship models. Manually feeding local memory back works but token costs climb fast.

Turning to the growing crop of "memory-as-a-product" startups, he finds no way to judge which is actually good: memory benchmarks use different setups, token budgets, and even different judge models, making scores incomparable. He asks the community how they choose. The post highlights a real gap in the agent memory ecosystem — no standardized, reproducible evaluation.

Original post →

More from coding & agent

coding & agent channel →