Tester: Unreleased Astra Solves Every Logic Puzzle Purely by Reasoning, vs 20-30% for sol 5.6
scaling01 · x · 2026-09-06
Blogger @FakePsyho reports testing the rumored (unverified) Astra model with a large batch of logic puzzles — and it solved all of them via pure logical reasoning, no code, including:
- World-championship-level puzzles
- Very large grids
- Unpublished puzzles
- Problems impossible to crack by backtracking alone
For comparison, sol 5.6 managed only 20-30% on similar tests. After analyzing the results, he found every solution legit, and says Astra now explains solving paths better than he can. Whether OpenAI RL'd logic puzzles to death doesn't matter to him — the results are real.
Related event: Rumored Astra Model Outperforms Sol 5.6 on Logic Puzzles(2 posts)→
More from Models
- Naval's bike analogy for SFT vs RL explains why DeepSeek R1 shocked the industry — McDonaghMatthew · 2026-09-07
- DeepSeek-R1 grew reasoning with pure RL, no SFT — and that's what changed everything — McDonaghMatthew · 2026-09-07
- OpenAI's next pretraining model rumored to be codenamed "Bel", unconfirmed by OpenAI — teortaxesTex · 2026-09-07
- Astra identifies sounds from mel spectrograms zero-shot, and users say we've barely scratched the surface — TomLikesRobots · 2026-09-07
- Next-gen model training could hit 5x pretraining compute if 300k GB200 rumor holds — scaling01 · 2026-09-07
- Banned triton is reward hacking; unbanned PyPI shortcut is a task bug — xeophon · 2026-09-07