GPT-6 Astra Sets New Records Across Benchmarks With More Intelligence Per Token
After rolling out to some users around Sept 5, GPT-6 Astra set new records across multiple third-party benchmarks, with the core narrative being "stronger and more token-efficient": Artificial Analysis's updated index shows it outscoring GPT-5.6 Sol while defining a new Pareto frontier of intelligence versus output tokens, cutting output tokens by roughly 10% at maximum reasoning; it scored a record 98.1 (xhigh) and 97.7 (high) on the Extended NYT Connections benchmark at about 40% lower cost per puzzle; and it ranked first on Terminal Bench 4.0 with the Codex harness at half the cost of the runner-up, roughly 47% below prior solutions.
Confirmed
- Artificial Analysis: GPT-6 Astra surpasses GPT-5.6 Sol and sets a new Pareto frontier on intelligence index versus output tokens (m2, m19)
- NYT Connections: 98.1 at xhigh, 97.7 at high, beating GPT-5.6 Sol at 40% lower cost per puzzle (m4, m18)
- Terminal Bench 4.0: first place at half the runner-up's cost, 47% cheaper than before (m7, m17)
- eyebench-v3: Astra-max scored 95% with 1/3.8 the output tokens and half the cost of Sol-max; adonissingh argues it leads on "intelligence per token" even at low reasoning effort (m12, m15, m16)
- MineBench.ai says Astra's first-generation output quality far exceeds the current top four models, the biggest quality jump since launch (m3)
- Hands-on: kimmonismus found Astra-Medium roughly matches GPT-5.6 xhigh at 1/3 the cost; Reddit user dubesor preferred its visual magazine layouts over GPT-5.6 Sol on iOS; Angaisb was satisfied at Max reasoning; Simon Willison's pelican SVG comparison across Sol, Terra and Luna peaked at 63 cents per run (m5, m6, m8, m9, m11)
- API pricing: input $10 / output $50 per 1M tokens, a 2.5x premium over GPT-5.6 Sol ($4 / $20) (m14)
Unconfirmed
- A FrontierMath T4 leak claims Astra scores above 5.6 Sol Pro (Max); commenter teortaxesTex doubts open-source models will match this within 18 months — personal opinion, no official data (m10)
- Developer xeophon's early take: incremental improvement rather than a step change — code quality similar and still needing steering, data analysis at the same level (m13)
- Most benchmark data comes from community and third-party leaderboards; OpenAI has not systematically confirmed it (m2, m3, m12)
Why it matters
- Token efficiency is seen as Astra's underdiscussed edge: Reddit user Cagnazzo82 argues that fewer tokens per task, combined with speed and quality, strikes directly at competitors' cost weakness (m1)
- DataLearnerAI notes the 2.5x API premium did not bring proportional reasoning gains — improvements concentrate in agentic tasks — so the "premium" and "token-saving" narratives coexist and real cost depends on use case (m14)
- If the "intelligence per token" advantage holds, it could reshape how model value is assessed and pressure rivals' pricing and product strategies
2026-09-05 ~ 2026-09-06 · 19 related posts
- Episode 1: GPT-6 Astra's 3D modeling demos flood in, community calls it SOTA(2026-09-04, 37 posts)
- Episode 2: Unverified 'GPT-6 Astra' Reportedly Generates Impressive 3D Models(2026-09-04, 2 posts)
- Episode 3: OpenAI Launches GPT-6 Astra and Declares the AGI Era(2026-09-04, 21 posts)
- Episode 4: GPT-6 Astra Sets New Records Across Benchmarks With More Intelligence Per Token(2026-09-05, 19 posts)
- Episode 5: GPT-6 Astra finds up to 176x code speedups in five minutes(2026-09-05, 2 posts)
- Episode 6: GPT-6 Astra floods social media with demos as official showcase draws criticism(2026-09-05, 9 posts)
- Episode 7: OpenAI's Astra Reportedly Trained on Over 100,000 GPUs(2026-09-05, 2 posts)
- Episode 8: GPT-6 Astra reportedly scores 95% on robot control at 43% of prior cost(2026-09-05, 6 posts)
- Episode 9: GPT-6 Astra Reportedly Beats Gemini in Vision Benchmarks(2026-09-05, 3 posts)
- Episode 10: OpenAI Launches 24-Hour GPT-6 Astra Demo Challenge(2026-09-05, 4 posts)
- Episode 11: GPT-6-Astra tops MathArena leaderboard(2026-09-06, 2 posts)
- Episode 12: GPT-6 Astra Tops Code Arena WebDev Leaderboard at 1797(2026-09-06, 4 posts)
- Episode 13: Leaked Benchmarks Claim GPT-6 Astra Aces Enterprise Tasks(2026-09-06, 2 posts)
- Episode 14: GPT-6 Astra Rebuilds Maya Ruins 3D Model in Two Hours(2026-09-06, 3 posts)
- Episode 15: Developer Says OpenAI's Astra Is First Model to Make Real Progress on His Ultra-Complex Project(2026-09-06, 2 posts)
- Episode 16: Rumor: GPT-6 Astra Beats Portal Autonomously, Doubts Remain(2026-09-06, 5 posts)
- Episode 17: GPT-6 Astra saturates spatial reasoning benchmarks(2026-09-06, 4 posts)
Primary sources
- GPT-6 Astra Beats 5.6 Sol Pro (Max) on FrontierMath T4; Open Models Seen 18 Months Behind — inductionheads · 2026-09-05
- Artificial Analysis: GPT-6 Astra cuts output tokens ~10% vs GPT-5.6 Sol at max effort — peterjliu · 2026-09-05
- GPT-6 Astra sets new NYT Connections benchmark record at 98.1 while costing ~40% less — zero0_one1 · 2026-09-05
- Early hands-on compares GPT-6 Astra vs GPT-5.6 Sol at max reasoning effort — Angaisb_ · 2026-09-05
- GPT-6 Astra's Untold Story May Be Token Efficiency, Argues Reddit User — Cagnazzo82 · 2026-09-05
- GPT-6 Astra tops Terminal Bench 4.0 at half the cost of #2 — charliermarsh · 2026-09-05
- MineBench Says GPT-6 Astra's First Generation Beats Current Top 4 Models — Ballist1cGamer · 2026-09-05
- GPT-6 Astra Max beats GPT-5.6 Sol at visual essays and magazine spreads, user finds — dubesor · 2026-09-05
- Simon Willison benchmarks GPT-6 Astra vs GPT-5.6 on SVG: max effort costs 63 cents per call — gaganghotra_ · 2026-09-05
- Early GPT-6 Astra tests: same intelligence as GPT-5.6 xhigh at roughly 1/3 the cost — charliermarsh · 2026-09-05
- [source] GPT-6 Astra runs Terminal Bench 4.0 at 47% lower cost — sherwinwu · 2026-09-05
- [source] Astra-max hits 95% on eyebench-v3 at half the cost of Sol-max, tokens ~3.8x fewer — adonis_singh · 2026-09-05
- Blogger: Astra-max's eyebench-v3 dominance is 'just not fair' — adonis_singh · 2026-09-05
- Astra-max claims vastly better intelligence-per-token even at low reasoning — adonis_singh · 2026-09-05
- [source] GPT-6-Astra beats GPT-5.6-Sol on updated Artificial Analysis Intelligence Index — DaserTheLaser · 2026-09-05
- GPT-6 Astra Costs 2.5× More Than GPT-5.6 Sol, but the Gains Are Agentic, Not Reasoning — DataLearnerAI · 2026-09-06
- Early Astra hands-on: code quality and data analysis feel incremental, says developer — xeophon · 2026-09-06