GPT-6 Astra Launches with Benchmark Leaks: ARC-AGI-3 Hits 98.6%
OpenAI 于 09-04 发布 GPT-6,代号 Astra,基准测试成绩随发布同步流出,多项指标大幅刷新纪录。据称这是 OpenAI 迄今规模最大的训练运行,也是首个在 Stargate 基础设施上用超过 10 万 DBU 预训练的模型。
已确认
- 多条帖子一致转述的基准成绩:ARC-AGI-3 从 7.8% 升至 98.6%,FrontierMath Tier 4 v2 达 97.6%,DeepSWE v1.1 达 74.1%,BenchCAD 95.9%,GPQA Diamond 96%,ExploitBench 满分,在展示的所有测试中全面超过 GPT-5.6 Sol(m4、m5、m11、m14 等)
- 官方博客口径:AAII 得分 61.2,略高于 GPT-5.6 Sol;Coding Agent Index 得 67,略落后于 Fable 5(m6、m17)
- 训练细节:研究员 Aidan Clark 称这是 OpenAI「迄今规模最大的训练运行」,首个在 Stargate 基础设施上以超 10 万 DBU 预训练的模型,且首次由前代模型主要参与监督下一代模型的训练(m10、m15、m20)
尚未确认
- 多数成绩以截图形式流传,部分帖子(m8、m13)明确标注未经官方证实,个别帖子提醒发布前跑分爆料常有夸大或伪造;FrontierMath Tier 4 得分在不同帖子中分别出现 97.6% 与 97.4% 两种说法,具体数值有待官方口径确认
为什么重要
- ARC-AGI-3 从 7.8% 跃升至 98.6% 的幅度极为罕见,randlongevity 表示完全没预测到该分数能这么高,认为这印证了「我们正活在奇点里」(m19)
- bindureddy 实测后给出更细的判断:Astra 在推理、数学、数据分析和研究能力上确实胜过 Fable 5.1,但 Fable 5.1 仍是编程之王(m12),与官方编码指数略逊 Fable 5 的数据互相印证
- teortaxesTex 借此感叹模型迭代节奏过快:GPT-5.6 Sol 发布不到两个月即被超越,「享乐跑步机」已达峰值,刚发布就过时(m7)
2026-09-04 ~ 2026-09-04 · 31 related posts
- Episode 1: Astra Rumored to Have Two Test Builds Focused on Long-Horizon Autonomy(2026-09-03, 2 posts)
- Episode 2: GPT-6 Astra Launches with Benchmark Leaks: ARC-AGI-3 Hits 98.6%(2026-09-04, 31 posts)
Primary sources
- OpenAI's Astra trained on largest-ever run using 100K+ GPUs at Texas Stargate site — cedric_chee · 2026-09-04
- GPT-6 Astra benchmarks: how the new model scores — CounterReady4774 · 2026-09-04
- GPT-6 Astra Benchmarks Released — james_6732 · 2026-09-04
- Leaked GPT-6-Astra benchmarks claim 98.6% on ARC-AGI-3, unverified — scaling01 · 2026-09-04
- Unconfirmed GPT-6 Astra benchmark screenshots circulate online — haider1 · 2026-09-04
- GPT-6 Astra hits 97.4% on FrontierMath Tier 4, trained on 100k+ GPUs — ChrisGPT · 2026-09-04
- Unconfirmed GPT-6 Astra benchmarks leak amid 'AGI era' launch claims — mark_k · 2026-09-04
- GPT-6 Astra's SWE results show only ~6% gain over its predecessor, critic notes — robleclerc · 2026-09-04
- GPT-6 Astra beats Fable 5.1 on Terminal science and Automation benchmarks — ChrisGPT · 2026-09-04
- Leaked OpenAI Astra scores claim 98.6% on ARC-AGI-3 and near-perfect math benchmarks — teortaxesTex · 2026-09-04
- [source] GPT-6 Astra full benchmarks: ARC-AGI-3 jumps from 7.8% to 98.6%, GPQA 96% — kimmonismus · 2026-09-04
- ARC-AGI-3 hits 98.6% with GPT-6, settling a five-month-old saturation bet — burny_tech · 2026-09-04
- GPT-6 Astra reportedly scores 98.6% on ARC-AGI-3, months ahead of Chollet's timeline — burny_tech · 2026-09-04
- Experts stunned as OpenAI's Astra posts a shocking ARC benchmark result — salahuddin · 2026-09-04
- Astra Is Out and So Are the Benchmarks — Hyleal · 2026-09-04
- [source] GPT-6 Astra benchmarks leak: 98.6% on ARC-AGI-3, biggest training run yet — Dr_Singularity · 2026-09-04
- More GPT-6 Astra training details: first model supervised by prior models — Dr_Singularity · 2026-09-04
- OpenAI's Astra beats Fable 5.1 on reasoning and math, but Fable 5.1 remains the coding king — bindureddy · 2026-09-04
- GPT-6 Astra reportedly scores 61 on Artificial Analysis Intelligence Index, prompting questions — thesaraharminta · 2026-09-04
- GPT-6 Astra Early Benchmarks Look Meh, as Release Cadence Hits Peak Burnout — teortaxesTex · 2026-09-04
- GPT-6 benchmarks leak online ahead of official release — thesaraharminta · 2026-09-04
- OpenAI pre-trained Astra on 100,000+ GPUs, its largest training run ever — haider1 · 2026-09-04
- GPT-6 Astra reportedly scores 61.2 on AAII, slightly above GPT-5.6 Sol — Angaisb_ · 2026-09-04
- 'We're living in the singularity': researcher stunned by ARC AGI 3 score — rand_longevity · 2026-09-04
- Leak claims GPT-6 Astra scores 98.6% on ARC-AGI-3 and tops most benchmarks — yuwen_lu_ · 2026-09-04
- "If These Evals Are True, We Need Far More People Working on Evals" — HarveenChadha · 2026-09-04
- Exploit Bench at 100%? New Model's Cybersecurity Score Stuns Developers — jarrodwatts · 2026-09-04
- [source] GPT-6 Astra blog reveals 61.2 on AAII, Coding Agent Index of 67 slightly behind Fable 5 — Angaisb_ · 2026-09-04
3 near-duplicate retellings: teropa · iamaliveix · burhop