GPT-6.1 Sol benchmarks: near-Astra performance at a fraction of the cost
After OpenAI released GPT-6.1 Sol, third-party evaluations poured in, with the general conclusion that its performance approaches the flagship GPT-6 Astra while costs drop substantially. Artificial Analysis's evaluation shows GPT-6.1 Sol pushes out OpenAI's cost-efficiency frontier overall: at Max effort, a single Intelligence Index task costs $0.72, below GPT-6 Astra's corresponding figure, and about 30% cheaper per task than GPT-6 Sol. Independent developers' hands-on tests largely agree.
Confirmed
- Pawel Huryn planted 105 bugs in 2 real code repositories for a blind test (Bug Hunt Bench, single prompt, scored across cost/time/turns): at max/xhigh effort, GPT-6.1 Sol fixed about 44 bugs for just $6.56, while GPT-6 Astra fixed 45 for $33; he had previously published results for all effort tiers (max/xhigh at n=3, others n=2) and argues that with cost factored in, there's no reason to keep using Astra
- adonissingh tested on eyebench-v3 and found GPT-6.1-Sol ranked second, with performance close to Astra at roughly 3.8x lower cost and about 1/8 the price of Opus-5.5; it's Pareto-optimal on cost, though output token efficiency still trails Astra across effort tiers
- cedricchee found 3D generation quality close to Astra and about 30% faster; codestantine said it beat Astra on his own low-resource language translation benchmark (reshared by nickbaumann)
- danielmckinn0n's GamowLabs RareBench results show GPT-6.1 Sol nearly matching Astra at about 1/4 the price
- Artificial Analysis has published evaluations for all GPT-6.1 Sol effort tiers (Low through Max) and simultaneously updated Intelligence Index v4.3 (replacing old items with AutomationBench-AA)
- Reddit user therealjerseytom said it "lives up to the hype": it handles the challenging work 6 Astra can do, but at significantly lower token cost, especially in scenarios requiring rework
Unconfirmed
- BridgeBench results show GPT-6.1 Sol only 3 points above GPT-6 Sol and 82 points behind GPT-6 Astra, a clear divergence from OpenAI's official "near-Astra intelligence" claim and other evaluations; the conclusion awaits more independent verification
- Wsz2020, citing AA data, argues the headline figure that "Sol max is 88% cheaper than Opus 5.5 max" is misleading (due to Opus's different pricing structure); direct cross-model price comparisons should be treated with caution
- The Runescape Bench reshared by banteg (a joke ranking) puts it second at roughly 10% of Astra's price, with limited reference value
- User chogku claims the GPT-6.1 Sol + Opus 5.5 combo approaches Astra-level quality at lower cost (reshared by rudrank), a personal opinion
Why it matters
Pawel Huryn stressed that unlike GPT-6 Sol, viewed as a "watered-down GPT-5.6 Terra," GPT-6.1 Sol is a genuine capability advance. If the low-cost, high-performance conclusion holds across more scenarios, it could significantly reshape the value-for-money landscape of frontier models, squeezing the usage of flagships like Astra in everyday coding and reasoning tasks.
2026-09-30 ~ 2026-10-02 · 21 related posts
- Episode 1: OpenAI Launches GPT-6 Sol and Luna at Half Price; Big Hallucination Gains but Not a Clean Sweep(2026-09-23, 92 posts)
- Episode 2: GPT-6 Sol/Luna launch splits opinion: half-price but dumber?(2026-09-23, 9 posts)
- Episode 3: LMArena Opens GPT-6 Sol and Luna for Public Testing(2026-09-23, 2 posts)
- Episode 4: Four New Models Reshape the Cost-Performance Frontier, Opus 5.5 Tops Intelligence Index(2026-09-23, 2 posts)
- Episode 5: OpenAI and Anthropic Slash Prices Within 90 Minutes, Kicking Off a New Model Price War(2026-09-23, 7 posts)
- Episode 6: Tests Show GPT-6 Sol Scales Nearly Linearly With Effort(2026-09-23, 3 posts)
- Episode 7: GPT-6 Astra Tops Bug Hunt Bench While Sol Version Regresses(2026-09-24, 2 posts)
- Episode 8: Leaks Map GPT-6 Codenames to GPT-5.6 Versions(2026-09-24, 3 posts)
- Episode 9: Bug Hunt Bench Ranks Frontier Coding Models on 105 Real Bugs with 200x Cost Gap(2026-09-24, 3 posts)
- Episode 10: GPT-6 Sol and Luna Launch With Aggressive API Pricing(2026-09-25, 2 posts)
- Episode 11: OpenAI's efficient GPT-6 models seen as freeing compute for bigger frontier models(2026-09-25, 2 posts)
- Episode 12: GPT-6 Sol Underestimated: 1.6x Robot Task Gains at 47% Lower Cost(2026-09-25, 2 posts)
- Episode 13: GPT-6.1 Sol Launches with Astra-Level Intelligence at a Fifth of the Price(2026-09-29, 32 posts)
- Episode 14: GPT-6 Astra Ultrafast Reportedly 8x Faster in Codex(2026-09-30, 2 posts)
- Episode 15: GPT-6.1 Sol benchmarks: near-Astra performance at a fraction of the cost(2026-09-30, 21 posts)
Primary sources
- GPT-6.1 Sol costs a quarter of GPT-6 Astra per task, pushing OpenAI's cost-efficiency frontier — ArtificialAnlys ·
- GPT-6.1 Sol nearly matches flagship on 105 planted bugs, at ~1/5 the cost — PawelHuryn ·
- Artificial Analysis posts full GPT-6.1 Sol evals, rolls out Intelligence Index v4.3 — ArtificialAnlys ·
- GPT-6.1 Sol near-Astra quality in 3D work while running ~30% faster — cedric_chee · 2026-09-30
- GPT-6.1-Sol Takes #2 on eyebench-v3, ~8x Cheaper Than Opus-5.5 — adonis_singh · 2026-09-30
- Cost Pareto-dominant, but second only to Astra on output-token efficiency — adonis_singh · 2026-09-30
- GPT 6.1 Sol Only +3 on BridgeBench, 82 Points Behind Astra: Benchmaxing Suspected — RexDouglass · 2026-09-30
- GPT-6.1 Sol scores 44 on real repos for just $6.56, near-top accuracy at 1/5 the cost — PawelHuryn · 2026-09-30
- GPT-6.1 Sol fixes 44 of 105 planted bugs for $6.56, matching Astra at a fraction of the cost — PawelHuryn · 2026-09-30
- GPT-6.1 Sol lands #2 on 'Runescape Bench' at ~10% the price of Astra — banteg · 2026-09-30
- Bug Hunt Bench updates all effort tiers; GPT-6.1 Sol xhigh nearly free on 105 planted bugs — PawelHuryn · 2026-09-30
- Early GPT-6.1 Sol user report: similar work done at dramatically lower token cost than 6 Astra — therealjerseytom · 2026-09-30
- GPT-6.1 Sol + Opus 5.5 combo called near-Astra quality at a fraction of the cost — rudrank · 2026-10-01
- GPT-6.1 Sol beats Astra on low-resource translation, dev benchmark shows — nickbaumann_ · 2026-10-01
- GPT-6.1 Sol benchmarks across all effort levels land, making GPT-6 Astra hard to justify — PawelHuryn · 2026-10-01
- GPT-6.1 Sol Benchmarked Across All Effort Levels; Atra Faster but Pricier — PawelHuryn · 2026-10-01
- GPT-6.1 Sol cuts cost per task ~30% further, now roughly a third of GPT-5.6 Sol — ArtificialAnlys · 2026-10-01
- [source] GPT-6.1 Sol costs a quarter of GPT-6 Astra per task, pushing OpenAI's cost-efficiency frontier — ArtificialAnlys · 2026-10-01
- [source] Artificial Analysis posts full GPT-6.1 Sol evals, rolls out Intelligence Index v4.3 — ArtificialAnlys · 2026-10-01
- Opus 5.5 high vs GPT-6.1 Sol max: AA data shows Sol's value edge is mostly latency, not intelligence — Wsz2020 · 2026-10-02
- GPT-6.1 Sol nearly matches Astra at a quarter of the price on RareBench — danielmckinn0n · 2026-10-02
3 near-duplicate retellings: cedric_chee · adonis_singh · PawelHuryn