Pokemon benchmark Paradigm 3: Astra generalizes to scrambled maps and fan-made games while rivals memorize
gleech · x · 2026-09-24
The Paradigm 3 team (Zanichelli, Leech, Grietzer) tested frontier models on Pokemon with a standardized, lightly-scaffolded harness to probe out-of-distribution generalization.
Key findings:
- Newer models are far better: Fable, Opus 5, Sol and Astra solve the tricky Rock Tunnel quickly and reliably where GPT-5.5 and Opus 4.8 usually failed.
- But evaluation is confounded: granular Pokemon training data saturates the web, and Pokemon-specific RL environments are plausible; newer models also recall game milestones better.
- Two out-of-distribution tests: on scrambled map variants, Sol and Opus 5 collapse while Astra doesn't drop at all; on the obscure fan game Pokemon Brown, Astra finished in 10k steps (5 human hours) with no notable memorization.
- Behavioral quirk: newer Claudes dutifully fight all battles; OpenAI models skip as many as possible.
Conclusion: evidence of memorization and shallow generalization in many models, but also reliability and tolerance to task variants (especially Astra) — signs of both increasing "general intelligence" and "whack-a-mole" task-specific post-training.
More from Models
- GPT-6 Astra only model to finish DrivingBench: 134.7m in 5:22 for $7.74 — petrusenko_max · 2026-09-24
- Opus 5.5 usage quota holds up well, user says it never ran out before reset — dotey · 2026-09-24
- Why GPT-6's rumored recurrent-depth architecture could change inference economics — panic_in_the_galaxy · 2026-09-24
- Jev, a No-Text Model Claiming 200x Faster Decisions, Sparks 'Is a Well-Formatted Wrong Answer Still a Hallucination?' — jamesbrooksco · 2026-09-24
- Hands-on: Jev Beats Gemini Flash on Accuracy, Latency, and Cost — cantrell · 2026-09-24
- User cancels switch to Codex after Opus 5.5's 40% price cut wins him back — ayushtweetshere · 2026-09-24