What learning from experience looks like: an agent learns to value draws in chess

anirudhg9119 · x · 2026-10-09

A concrete case from the paper: in Chess, the agent scores 0 at the 300-ply limit, then changes its engine to value draws near the limit. At checkpoint 18 it holds a draw until Stockfish errs, then mates—a persistent fix learned from experience. The NetHack example is covered in the thread's main entry.

Related event: Agent Plasticity: A New Metric for Measuring How Agents Self-Improve Through Experience(13 posts)→

Original post →

More from Research

Research channel →