Yoav Goldberg: Model Excels at Sokoban-Like Puzzles—Trained on Them?
yoavgo · x · 2026-09-26
Yoav Goldberg notes frontier agents are remarkably good at solving grid-based, turn-based sokoban-like puzzle games and asks if they were trained on such games. Lukasz Kaiser responds that strength emerges wherever there's enough diverse task data plus a good grader (arXiv lemmas as an example), and from using the agents it seems they're generalizing fairly well rather than memorizing.
Related event: Frontier Model Excels at Sokoban Puzzles, Raising Training Data Questions(2 posts)→
More from Models
- Observation: Astra uses filler tokens far more effectively than other models — scaling01 · 2026-09-26
- Ethan Mollick: 'Keep prompts short' is bad advice, and minimizing token cost confuses inputs with outputs — emollick · 2026-09-26
- GPT-6 Luna uses fewer reasoning tokens than 5.6 on ARC-AGI-2, hard tasks stymie both — mhmazur · 2026-09-26
- Heavy user: fast, cheap Claude Opus 5.5 now takes all my serious work — brandon_galang · 2026-09-26
- Model excels at Sokoban-style puzzles, sparking questions about training data contamination — lukaszkaiser · 2026-09-26
- GPT-6 Astra's First Draft Fooled Every's CEO: Big Writing Upgrade, But Some Bad Habits — every · 2026-09-26