GPT-5.6 Fails at Tracking Board Game States in Images
Hot_Path4185 · reddit · 2026-07-17
A Reddit user tested GPT-5.6 (High) on its ability to track board game states across a sequence of images, finding its performance unstable in Nine Men's Morris.
The testing method involved:
- Using screenshots from the mini-game in Assassin's Creed IV: Black Flag
- Providing a new image of the board state for every move
- Asking the model to continue playing based on the sequence of images
Results showed that the model gradually lost memory of piece positions. It even made moves that ignored existing pieces on the board, ultimately failing to find a clear winning path.
The poster wanted to know if this is a known weakness: whether the model struggles specifically with visual board state tracking, or if it's a broader issue with sequential image reasoning.
More from Models
- OpenAI rolls out voice in GPT-Live, but the UI obscures search and reasoning — Graham_dePenros · 2026-07-22
- Gemini 3.6 Flash goes live in Antigravity with 17% fewer output tokens — rseroter · 2026-07-22
- Moonshot’s Kimi K3 sets a new open-weights ECI record at 156 — scaling01 · 2026-07-22
- Nanbeige4.2-3B launches as a 3B Looped Transformer model that beats larger baselines — Wooden-Deer-1276 · 2026-07-22
- A post says six companies now beat Google’s best LLM, including two open-source models — soham_btw · 2026-07-22
- Gemini 3.6 Flash benchmark results reignite concerns that Google is slipping behind — minxio_ · 2026-07-22