Raschka: looped transformers aren't why GPT-6 Astra's reasoning is less monitorable
机器之心 · wechat · 2026-09-14
Sebastian Raschka's deep dive into GPT-6 Astra: benchmarks (99.9% on ARC-AGI-3 vs 7.8% for GPT-5.6 Sol), standout computer-use capability trained in macOS RL environments built on tens of thousands of Macs, 100k Grace Blackwell GPUs, and a full explanation of looped transformers from Universal Transformer to Nanbeige4.2-3B and Mixture-of-Recursions. His verdict: recurrent depth is not the cause of reduced chain-of-thought monitorability — stronger models simply backtrack less, as shown by Luna needing 80% more tokens than Sol for similar performance.
Related event: Raschka Deep-Dives GPT-6 Astra's Looped Transformer Architecture(2 posts)→
More from Models
- Leaker teases 'more exciting OpenAI releases this week' at DevDay-level ship volume — MickeySteamboat · 2026-09-15
- Mistral audio research lead: voice AI needs a screen, won't replace it — Machine Learning Street Talk · 2026-09-15
- User argues uneven AI subscription tiers: $200 plan gets 20× usage vs 5× at $100 — NoOne_n13 · 2026-09-15
- OpenAI reportedly pauses $200 ChatGPT Pro signups as GPT-6 Astra demand saturates capacity — emmanuelvivier · 2026-09-15
- Vibe coding 3D game dev: Claude Pro beats ChatGPT Plus on limits and code audits — AdvertisingBubbly546 · 2026-09-15
- TerminalBench scoring isn't comparable: Astra gets zeros on safety stops, Fable re-routes to Opus — xeophon · 2026-09-15