Raschka: looped transformers aren't why GPT-6 Astra's reasoning is less monitorable

机器之心 · wechat · 2026-09-14

Sebastian Raschka's deep dive into GPT-6 Astra: benchmarks (99.9% on ARC-AGI-3 vs 7.8% for GPT-5.6 Sol), standout computer-use capability trained in macOS RL environments built on tens of thousands of Macs, 100k Grace Blackwell GPUs, and a full explanation of looped transformers from Universal Transformer to Nanbeige4.2-3B and Mixture-of-Recursions. His verdict: recurrent depth is not the cause of reduced chain-of-thought monitorability — stronger models simply backtrack less, as shown by Luna needing 80% more tokens than Sol for similar performance.

Related event: Raschka Deep-Dives GPT-6 Astra's Looped Transformer Architecture(2 posts)→

Original post →

More from Models

Models channel →