324 chess games show LLM self-notes act as memory but can't fix board blindness

marvijo-software · reddit · 2026-10-12

An experiment had GPT-6.1 Sol, Claude Sonnet 5.5 and DeepSeek V4.1 Flash play 324 chess games across 31 tournaments, writing and re-reading notes between games. No upward trend emerged: the first blunder was almost always an undefended piece around move 13-16, notes never prevented it, models stuffed notes with opening lines, and board images helped less than text. Conclusion: notes work as memory, not board vision. Notes are open-sourced.

Original post →

More from coding & agent

coding & agent channel →