Raschka deep dive: GPT-6 Astra, recurrent depth, and whether it hides its chain of thought
bibryam · x · 2026-09-09
Sebastian Raschka published a detailed article on GPT-6 Astra and the rumor that it 'hides' its reasoning trace, tying it to looped transformers and recurrent depth.
- Impressions: Astra is the best model he has used, leapfrogging GPT-5.6 across writing, math and coding, with especially strong 3D rendering/animation demos; benchmark results back this up
- He explains what looped transformers are (reusing transformer blocks) and how they relate to hidden chains of thought
- Surveys new insights from recent research papers on looping transformer blocks
A solid primer for anyone following the recurrent-depth architecture direction.
Related event: Raschka Deep-Dives into GPT-6 Astra and Looped Transformers(3 posts)→
More from Models
- Thomson Reuters' 397B Qwen-based legal LLM beats GPT-5.4 on domain factuality — DeepLearningAI · 2026-09-09
- 27B 1-bit model runs in the browser at 25-30 tok/s on a 6GB RTX 3060 laptop (WebGPU) — mentria-ai · 2026-09-09
- GPT-6 Astra Pro/Ultra review: first LLM to nail one-shot prompts nearly every time — rschu · 2026-09-09
- Raschka dissects GPT-6 Astra: looped transformers, recurrent depth, hidden CoT — rasbt · 2026-09-09
- Agentic video's real story isn't watching 90 minutes—it's reasoning about where to spend compute — romitheguru · 2026-09-09
- GPT-6 Astra system card draws fire: OpenAI claims 'most aligned model' ever — TheZvi · 2026-09-09