LLM Memory Benchmarks Are Broken: ECC Tests Efficiency, Compression, Calibration
cindy_x_wu · x · 2026-09-02
Shmuel Berman points out a flaw in current LLM/VLM memory evaluation: a system can score 100% accuracy simply by re-reading the input for each question — behavior we'd call terrible memory in a human.
Accuracy alone doesn't measure real memory. The team proposes ECC, a framework that goes beyond accuracy along three axes: Efficiency, Compression, and Calibration — addressing the rereading loophole and giving memory-improvement research a fairer yardstick.
More from Models
- Gemini 3.8 Flash hits 59 on AA Intelligence Index, closing on open frontier at 60 — cedric_chee · 2026-09-03
- Fable 5.1 day one: more human-like, frontend still beats GPT models — Yuchenj_UW · 2026-09-03
- Matt Shumer: AI release slowdown is safety review bottleneck, dam about to break — mattshumer_ · 2026-09-03
- Fable 5.1 Touted as New #1 LLM, Holding Up in Early Coding Tests — davidthesong · 2026-09-03
- Polymarket puts 78% odds on OpenAI's rumored Astra model launching tomorrow — Polymarket · 2026-09-03
- Polymarket puts 78% odds on OpenAI's rumored Astra model shipping tomorrow — Polymarket · 2026-09-03