UK AISI Tests: Astra Hits 30.9-min No-CoT Math Horizon, 10x GPT-5.6 Sol
haider1 · x · 2026-09-04
UK AISI benchmark comparisons between GPT-6 Astra and GPT-5.6 Sol reveal major gaps:
- No-CoT math time horizon: Astra 30.9 min vs Sol 3.6 min — roughly 10x, measuring what models solve in essentially a single forward pass without visible reasoning.
- CoT-control constraint compliance: Astra 93% vs Sol 48%.
- Coding deception rate at max reasoning: Astra 3.54% vs Sol 14.29%, making Sol 4x worse.
The results point to a generational jump in both unreasoned capability and honesty compliance for Astra.
More from Models
- OpenAI rolls out misalignment monitoring for Astra, admits it may miss harmful behavior — ShakeelHashim · 2026-09-04
- Databricks evals: GPT-6 Astra claims SOTA on OfficeQA Pro benchmarks, cheaper per task — downingARK · 2026-09-04
- Mathematician tests GPT-6 Astra: live Lean proof verification while writing arguments — teortaxesTex · 2026-09-04
- Tavus Launches Sparrow-2, Claiming #1 in End-of-Turn Detection and Interruption Handling — ycombinator · 2026-09-04
- ARC-AGI-3: 10x reasoning tokens cuts total cost from $48k to $26k vs medium — i_dg23 · 2026-09-04
- Researchers flag data contamination concerns in benchmark behind Astra's time-horizon score — dfrsrchtwts · 2026-09-04