Agent Plasticity paper measures how efficiently AI agents turn experience into self-improvement
anirudhg9119 · x · 2026-10-09
The arXiv paper "Agent Plasticity: Measuring Self-Improvement Through Experience" argues existing benchmarks only measure what an agent can do at a fixed point, not how well it learns. Agents amortize experience into reusable artifacts inherited by future instances; performance is measured at checkpoints on training and held-out interactions, accounting for learning cost. Agent plasticity = held-out gain per unit of learning cost.
Key findings:
- Frontier models with identical learning opportunities show sharply different improvement trajectories: some gain persistent improvements, others stay near or below initial performance.
- In-regime gains often transfer only partially out-of-distribution; endpoint capability and acquisition efficiency diverge (e.g., one model reaches the highest fitted performance in Chess/Go/Hex while another is the most efficient learner).
- Artifact reuse correlates with gains but is insufficient: even agents reusing artifacts in 94–98% of relevant decisions improve by very different amounts, and strong improvers see 83–99% of remaining failures occur while an artifact is in use.
Concrete cases: in Chess, a Claude model scoring 0 at the 300-ply limit learns to value draws near the limit, later holding a draw until Stockfish errs and then mating; in NetHack, an agent dies praying via raw keystrokes, then adds a reusable rule to pray through its controller. Learning from experience is measurable but far from automatic.
More from coding & agent
- Open-source Android app Angel runs MCP servers on-device and hands tools to local or cloud models — ihaveaboyfriendsorry · 2026-10-09
- ClickUp's Brain² agent builds reports and dashboards with E2B microVM sandboxes — mathemagic1an · 2026-10-09
- TensorFold joins NVIDIA Inception, gets early access to next Nemotron for 0-day support — HankYeomans · 2026-10-09
- Strata rewrote its Git history to wipe all 'Co-Authored by Claude' evidence — dasbin · 2026-10-09
- Agents on both sides of Zapier and Retell AI sorted out a call-messaging webhook — ramagetime · 2026-10-09
- Building RL environments in 2026: 10% writing tasks, 90% preventing agent cheating — geoffwolfe · 2026-10-09