Yoav Goldberg: Model Excels at Sokoban-Like Puzzles—Trained on Them?

yoavgo · x · 2026-09-26

Yoav Goldberg notes frontier agents are remarkably good at solving grid-based, turn-based sokoban-like puzzle games and asks if they were trained on such games. Lukasz Kaiser responds that strength emerges wherever there's enough diverse task data plus a good grader (arXiv lemmas as an example), and from using the agents it seems they're generalizing fairly well rather than memorizing.

Related event: Frontier Model Excels at Sokoban Puzzles, Raising Training Data Questions(2 posts)→

Original post →

More from Models

Models channel →