Continuous learning benchmark has models learn chess over 200 games — Elo barely improves
imjustnewatai · x · 2026-10-04
Peter Gostev built a "continuous learning" benchmark where models are given the goal of learning to play chess against a Stockfish opponent over 200 games — they can pick difficulty and take notes, but no cheating via engines. A live site tracks their Elo journey.
Early results:
- GPT-6 Astra: finished its 200 games with a negative Elo trend (vs full-strength Stockfish: 66 losses, 2 draws in 68 games, rolling score 0.01)
- Opus: a slight positive improvement, possibly within random noise; still playing
The benchmark probes whether models can self-improve without a real continuous-learning mechanism — an early proxy for RSI. The retweeter says he wants to build something similar.
More from Models
- whurley rips into Google: Gemini can't even access all your Google accounts — whurley · 2026-10-04
- StepFun's Step 5 Preview debuts at #7 among open-weight models on Vals — StepFun_ai · 2026-10-04
- Fable 5.1 replaces Opus 5.5 as default recommended model — LillyPlayer · 2026-10-04
- Simon Willison's September newsletter: Fable-class models, a pricing war, LLMs for math — Simon Willison · 2026-10-04
- Opus 5.5 on Max 20x turbo now beats GPT Pro 200, as Codex desktop app decays — mertdumenci · 2026-10-04
- Memento work on context management heads to COLM, featuring a mementified Qwen3-32B — DimitrisPapail · 2026-10-04