UK AISI eval: GPT-6 Astra hits 30.9-min no-CoT math time horizon, 8x GPT 5.6 Sol
teortaxesTex · x · 2026-09-04
The UK AI Safety Institute published its evaluation of GPT-6 Astra:
- No-CoT math time horizon: 30.9 minutes vs 3.6 for GPT 5.6 Sol — Astra solves far harder math in a single forward pass.
- CoT controllability: follows constraints on 93% of samples vs 48% for GPT 5.6 Sol, though controllability still degrades over longer reasoning stretches.
- CoT legibility: Astra reasons in a more compressed style; raw reasoning is generally understandable but ambiguous phrases appear more often.
Related event: UK AISI Tests GPT-6 Astra: 30.9-Minute No-CoT Math Horizon(2 posts)→
More from Models
- Tester: AI-text detector Pangram shows zero false positives, but adversarial rewriting evades it — alex_peys · 2026-09-04
- GPT-6 Astra jumps 42.2pp to top Terminal-Bench-Science at 64.6% — BenBlaiszik · 2026-09-04
- OpenAI researcher teases upcoming model: 'best in class' at computer use, research, coding — athyuttamre · 2026-09-04
- FrontierMath is 'dead' after 664 days as frontier models conquer the math benchmark — willdepue · 2026-09-04
- Matt Shumer's GPT-6 Astra review: Manager Loop sustains a week-long Manhattan build in Unreal — mattshumer_ · 2026-09-04
- Screenshots of Amazon Astra's Chain-of-Thought Output Circulate on Reddit — Tough_North7059 · 2026-09-04