What learning from experience looks like: an agent learns to value draws in chess
anirudhg9119 · x · 2026-10-09
A concrete case from the paper: in Chess, the agent scores 0 at the 300-ply limit, then changes its engine to value draws near the limit. At checkpoint 18 it holds a draw until Stockfish errs, then mates—a persistent fix learned from experience. The NetHack example is covered in the thread's main entry.
More from Research
- Lancet study: patient-facing conversational AI holds up in real urgent care settings — EricTopol · 2026-10-09
- First Workshop on Agent Behavior at COLM 2026 Set for Oct 9 in San Francisco — _Hao_Zhu · 2026-10-09
- Autorubric ships 25-recipe cookbook for rubric design, judge calibration and cost control — deliprao · 2026-10-09
- Autorubric at COLM 2026: a unifying framework for rubric-based LLM evaluation — deliprao · 2026-10-09
- Google's AMIE medical AI publishes first prospective real-patient study in The Lancet — ymatias · 2026-10-09
- TraceExtract open-sourced: data engine for µ0 world model trained on video with zero action labels — RexDouglass · 2026-10-09