324 chess games show LLM self-notes act as memory but can't fix board blindness
marvijo-software · reddit · 2026-10-12
An experiment had GPT-6.1 Sol, Claude Sonnet 5.5 and DeepSeek V4.1 Flash play 324 chess games across 31 tournaments, writing and re-reading notes between games. No upward trend emerged: the first blunder was almost always an undefended piece around move 13-16, notes never prevented it, models stuffed notes with opening lines, and board images helped less than text. Conclusion: notes work as memory, not board vision. Notes are open-sourced.
More from coding & agent
- TikTok turns on vibe-coded Photoshop alternative; ex-Adobe dev says no one knows what's in Adobe's code either — jasonkneen · 2026-10-12
- Codex roadmap leak shows agent message board among 59 features in development — gajesh · 2026-10-12
- Claude Code creator: Opus 5.5 does a month of work in a day — and your old prompts now backfire — VeryWellVersed · 2026-10-12
- Solo dev ships Scape: Claude Code executes, Codex reviews in adversarial loop — croovies · 2026-10-12
- banteg builds a DSL to describe all 50 original game quests — banteg · 2026-10-12
- Bittensor SN33 claims skills turn small models into frontier-level performers — markjeffrey · 2026-10-12