Agent Plasticity: A New Metric for Measuring How Agents Self-Improve Through Experience
Harman Singh, Anuj Mahajan, Jason, and colleagues released the paper "Agent Plasticity: Measuring Self-Improvement Through Experience" (arXiv 2610.08902); the team also includes well-known researchers such as Rob Fergus and Sanjeev Arora. The paper's core claim: evaluating agents should consider not just what they can do, but how they improve, and it proposes the Agent Plasticity metric.
Confirmed
- Metric definition: the held-out performance gain per unit of learning cost, measured after the agent amortizes experience into reusable artifacts, up until saturation.
- Experiments were run on Chess, Go, and Hex; results show GPT-5.6 Sol's learning efficiency is roughly 5x that of Claude Fable 5.
- One key finding: the strongest agents aren't necessarily the best learners; even though all four models reuse existing artifacts in 94–98% of relevant decisions, gains still vary widely, meaning reuse alone can't explain learning outcomes.
- Case studies: @anirudhg9119 recounted several concrete examples in a thread. In chess, after a game ended 0 points at the 300-ply cap, Claude Fable 5 modified its own engine to value draws, and later checkpoints even held out against Stockfish until being checkmated. In NetHack, Claude Opus 5.5 once died by praying via raw keypresses, then added a reusable rule to its controller to pray through the controller instead—and that prayer later actually healed it, demonstrating persistent self-improvement.
Why it matters
- As @anirudhg9119 put it, "having memory isn't the same as learning": most agents struggle to improve from experience, and this work offers a quantifiable framework for distinguishing memorization from genuine self-improvement.
- A clue to why some agents barely improve is that they don't reuse what they've learned—but the data shows high reuse rates don't guarantee gains either, leaving the mechanisms of agent learning an open research question.
2026-10-09 ~ 2026-10-09 · 7 related posts
Primary sources
- Agent Plasticity paper proposes benchmarking how efficiently AI agents learn from experience — anirudhg9119 ·
- Agent Plasticity paper released with all-star author list incl. Rob Fergus, Sanjeev Arora — anirudhg9119 ·
- Agent Plasticity metric: GPT-5.6 Sol learns ~5x more efficiently than Claude Fable 5 — anirudhg9119 ·
- Agent Plasticity: top-performing agents aren't the most efficient learners — RulinShao · 2026-10-09
- [source] Agent Plasticity metric: GPT-5.6 Sol learns ~5x more efficiently than Claude Fable 5 — anirudhg9119 · 2026-10-09
- All four models reuse artifacts 94-98% of the time, yet their learning gains diverge widely — anirudhg9119 · 2026-10-09
- Agent Plasticity paper shows Claude learning reusable rules from NetHack and Chess failures — anirudhg9119 · 2026-10-09
- Persistent Self-Improvement: Why Having Memory Isn't the Same as Learning — anirudhg9119 · 2026-10-09
- [source] Agent Plasticity paper released with all-star author list incl. Rob Fergus, Sanjeev Arora — anirudhg9119 · 2026-10-09
- [source] Agent Plasticity paper proposes benchmarking how efficiently AI agents learn from experience — anirudhg9119 · 2026-10-09