LLM Memory Benchmarks Are Broken: ECC Tests Efficiency, Compression, Calibration

cindy_x_wu · x · 2026-09-02

Shmuel Berman points out a flaw in current LLM/VLM memory evaluation: a system can score 100% accuracy simply by re-reading the input for each question — behavior we'd call terrible memory in a human.

Accuracy alone doesn't measure real memory. The team proposes ECC, a framework that goes beyond accuracy along three axes: Efficiency, Compression, and Calibration — addressing the rereading loophole and giving memory-improvement research a fairer yardstick.

Original post →

More from Models

Models channel →