Agent Plasticity paper shows Claude learning reusable rules from NetHack and Chess failures
anirudhg9119 · x · 2026-10-09
The author shares examples of persistent self-improvement in agents from their paper:
- In NetHack, Claude Opus 5.5 died while praying via raw keystrokes, then added a reusable rule to pray through its controller instead — a later prayer actually healed it.
- In Chess, after a game scored 0 at the 300-ply limit, Claude Fable 5 modified its engine to value draws near the limit; at a later checkpoint it held a draw until Stockfish eventually mated.
The author frames these as fixes learned from experience rather than pre-programmed behavior.
More from Research
- Leak hints OpenAI runs multi-agent swarms on time budgets, not token budgets, and treats it as IP — maksym_andr · 2026-10-09
- 100+ Mathematicians React: OpenAI's Quasi-Riemann Result "Almost Unbelievable" — littmath · 2026-10-09
- Epoch launches Automation Reports: Claude Fable 5.1 and GPT-6 Astra lead but can't automate its research — scaling01 · 2026-10-09
- Iris-3B: pixel-space diffusion offers no edge over latent models, study finds — speridlabs · 2026-10-09
- Meta's MIMESIS: 9B user simulator beats GPT-5.5 for training interactive agents — facebook · 2026-10-09
- LightOn's Amélie Chatelain to teach free lesson on scaling late-interaction search, indexes 5-30x smaller — IgorCarron · 2026-10-09