Raschka on GPT-6 Astra: Looped Transformers, Computer Use, and Hidden CoT Rumors
Ahead of AI (Sebastian Raschka) · rss · 2026-09-09
Sebastian Raschka's deep dive on GPT-6 Astra: it's the best model he's used, especially at 3D rendering/computer-use tasks (99.9% on ARC-AGI-3 vs 7.8% for GPT-5.6 Sol), though independent agentic benchmarks show a narrower lead. He explains harness effects on evals, OpenAI's Mac fleet for computer-use RL environments, suggests pruning stale AGENTS.md/SKILL.md files, and walks through looped transformer research and its relation to rumors that Astra hides its chain of thought.
More from Models
- GPT-6 Astra tops RSI-Exam at 0.5126, 18.4% above GPT-5.6 Sol — HuaxiuYaoML · 2026-09-09
- APEX-Agents 1.1 benchmark update: Claude Fable 5.1 tops leaderboard at 68.6% — amaarora · 2026-09-09
- Frontier labs must shrinkflate the $200 subscription to upsell you to API pricing — StewartalsopIII · 2026-09-09
- User math: peak pricing 2x but base rate better, 252M cache tokens cost just $0.75/day — teortaxesTex · 2026-09-09
- Astra noticeably worse than Sol in long threads, dev finds handoff workaround — jdjohnson · 2026-09-09
- VC communism is over: frontier models on rationing force hard model-choice thinking — StewartalsopIII · 2026-09-09