FULL STORY

GPT-6 Astra: From Leak to Launch in Two Days

After a leaker revealed Astra's dual checkpoints and autonomous-task focus, OpenAI launched GPT-6 Astra the next day with record benchmark scores, officially positioning it as its most aligned model yet.

2026-09-03 ~ 2026-09-04 · 3 episodes · 46 posts

Episode 1 · Astra Rumored to Have Two Test Builds Focused on Long-Horizon Autonomy (2026-09-03, 3 posts)

Insider Lentils80 says two GPT Astra checkpoints, ultima-alpha and vega-alpha, are being tested with a focus on long-running autonomous tasks, while others call it a watershed model that breaks OpenAI's internal research benchmarks.

Episode 2 · GPT-6 Astra Launches with Benchmark Leaks: ARC-AGI-3 Hits 98.6% (2026-09-04, 41 posts)

OpenAI released GPT-6, codenamed Astra, on 09-04, with benchmark scores leaking alongside the launch that shattered records across multiple metrics. It is reportedly OpenAI's largest training run to date and the first model pretrained on Stargate infrastructure using more than 100,000 DBUs.

Confirmed

  • Benchmark figures consistently relayed across multiple posts: ARC-AGI-3 up from 7.8% to 98.6%, FrontierMath Tier 4 v2 at 97.6%, DeepSWE v1.1 at 74.1%, BenchCAD 95.9%, GPQA Diamond 96%, a perfect score on ExploitBench, and across all displayed tests a comprehensive win over GPT-5.6 Sol (m4, m5, m11, m14, etc.)
  • Official blog figures: AAII score of 61.2, slightly above GPT-5.6 Sol; Coding Agent Index of 67, slightly behind Fable 5 (m6, m17)
  • Training details: researcher Aidan Clark says this is OpenAI's "largest training run to date," the first model pretrained on Stargate infrastructure with over 100,000 DBUs, and the first time a previous-generation model primarily supervised the training of the next generation (m10, m15, m20)

Unconfirmed

  • Most scores circulated as screenshots; some posts (m8, m13) explicitly flag them as unofficial, and a few remind that pre-launch benchmark leaks are often exaggerated or fabricated; the FrontierMath Tier 4 score appears as both 97.6% and 97.4% in different posts—exact value awaits official confirmation

Why it matters

  • The jump from 7.8% to 98.6% on ARC-AGI-3 is extraordinarily rare; randlongevity said they never predicted the score could be this high, seeing it as proof that "we're living in the singularity" (m19)
  • After hands-on testing, bindureddy offered a finer judgment: Astra indeed beats Fable 5.1 on reasoning, math, data analysis, and research, but Fable 5.1 remains the king of coding (m12)—consistent with the official coding index slightly trailing Fable 5
  • teortaxesTex lamented the breakneck iteration pace: GPT-5.6 Sol was surpassed less than two months after launch—the "hedonic treadmill" has peaked, obsolete upon release (m7)

21 more related posts →

Episode 3 · OpenAI Calls Astra Its Most Aligned, Best Coding Model (2026-09-04, 2 posts)

OpenAI officially states GPT-6 Astra is its most aligned model with substantially improved intent understanding, and the best yet at computer use, software engineering, and vision.