Memory beats model choice: agents with memory hit 80% vs 45% for long context
alex_verem · x · 2026-08-18
Key takeaways from a four-year survey of agent memory (Generative Agents through MemoryArena):
- In the Generative Agents experiment, removing the reflection component made characters that had run coherent multi-day plans collapse into repetitive nonsense within days.
- On MemoryArena's multi-session tasks, agents with active memory completed over 80%, while the same setups leaning on a long context window alone dropped to roughly 45%.
- A bigger context window is not memory — reading everything and knowing what to recall are different skills.
- The paper's closing point: teams spend months benchmarking which model to use and give the memory architecture an afternoon, exactly backwards.
More from coding & agent
- Building a GTM machine with agents: from 0 to $10k MRR for one-person businesses — EXM7777 · 2026-08-18
- Sentence Transformers v6.0 Ships Late Interaction Models, Its Largest Update Yet — tomaarsen · 2026-08-18
- Breaking Changes in Sentence Transformers v6.0: transformers v5 Floor, API Shifts — tomaarsen · 2026-08-18
- bfloat16 Sigmoid Saturates CrossEncoder Rankings; Fix Lifts nDCG From 0.18 to 0.68 — tomaarsen · 2026-08-18
- From ModernBERT-base to 0.48 NanoBEIR nDCG@10 in 25 Minutes on One RTX 3090 — tomaarsen · 2026-08-18
- v6.0's Interpretability Module Renders Exact ColPali-Style Heatmaps From MaxSim — tomaarsen · 2026-08-18