Memory is Key for Long-Horizon Agents: 15 LLMs Manage a Football Club for 20 Years
dair_ai · x · 2026-08-30
A study tested long-horizon memory and planning by having LLM agents manage a football club for 20 in-game seasons (approx. 340-400 decision stops).
Key Findings:
- All 15 frontier models survived the entire horizon, while most scripted baselines failed.
- Scale, price, vendor, or token spend did not predict rankings; order only settled late in the run.
- Top models exhibited specific managerial behaviors: cutting slow-payoff investments near the end and opening contract renewals well before deadlines.
- Two universal failures were found: no model learned the market's hidden prices from rejected bids, and self-managed memory also failed.
The research highlights that memory/recall is critical for breaking long-horizon agent bottlenecks.
More from Research
- Anthropic shows AI researchers autonomously improving alignment of other models — VraserX · 2026-08-30
- Learn Positional Encodings derivation from first principles — zainhas · 2026-08-30
- COLM Paper Traces Capability Provenance in LLMs via Gradient Attribution — ziv_ravid · 2026-08-30
- Toby Ord paper argues recursive self-improvement has physical limits — Exponential View (Azeem Azhar) · 2026-08-30
- AI Formalization Tools Fable and Sol Spot First Repairable Error in Published Literature — Sauers_ · 2026-08-30
- Mark Schmidt Posts ICML Tutorial Video: Is Numerical Optimization Theory Irrelevant to ML Practice in 2026? — MarkSchmidtUBC · 2026-08-30