Continuous learning benchmark has models learn chess over 200 games — Elo barely improves

imjustnewatai · x · 2026-10-04

Peter Gostev built a "continuous learning" benchmark where models are given the goal of learning to play chess against a Stockfish opponent over 200 games — they can pick difficulty and take notes, but no cheating via engines. A live site tracks their Elo journey.

Early results:

The benchmark probes whether models can self-improve without a real continuous-learning mechanism — an early proxy for RSI. The retweeter says he wants to build something similar.

Original post →

More from Models

Models channel →