Pokemon benchmark Paradigm 3: Astra generalizes to scrambled maps and fan-made games while rivals memorize

gleech · x · 2026-09-24

The Paradigm 3 team (Zanichelli, Leech, Grietzer) tested frontier models on Pokemon with a standardized, lightly-scaffolded harness to probe out-of-distribution generalization.

Key findings:

Conclusion: evidence of memorization and shallow generalization in many models, but also reliability and tolerance to task variants (especially Astra) — signs of both increasing "general intelligence" and "whack-a-mole" task-specific post-training.

Original post →

More from Models

Models channel →