FULL STORY
GPT-6.1 Sol: Strong Independent Tests, Contested Internal Benchmarks
After GPT-6.1 Sol's launch, independent testers found performance near GPT-6 Astra at far lower cost. OpenAI's internal benchmarks claiming a big lead over Claude Opus 5.5 soon drew community skepticism.
2026-09-30 ~ 2026-10-01 · 2 episodes · 18 posts
Episode 1 · GPT-6.1 Sol benchmarks land: near-Astra performance at a fraction of the cost (2026-09-30, 16 posts)
Since GPT-6.1 Sol launched, multiple independent reviewers have tested it on coding, 3D generation, translation and other tasks. The mainstream verdict: performance close to GPT-6 Astra with dramatically lower token costs, giving it a clear edge in value for money—though a few benchmark results diverge from the official "near-Astra level" claim.
Confirmed
- Pawel Huryn ran a blind test (Bug Hunt Bench) planting 105 bugs across 2 real code repositories: at max effort, GPT-6 Astra fixed 45 for $33, while GPT-6.1 Sol fixed 44 for just $6.56—roughly 1/5 of Astra's cost or less. He judged this one "the real deal," unlike GPT-6 Sol, which was seen as a watered-down GPT-5.6 Terra.
- Huryn later published scores for all effort tiers (n=3 for max/xhigh, n=2 for the rest, with follow-up reruns ongoing), concluding that Astra is faster but pricier—given the costs, he sees no reason to keep using GPT-6 Astra.
- adonissingh's eyebench-v3 leaderboard ranks GPT-6.1-Sol second, displacing Opus-5.5; it's about 3.8x cheaper than astra and roughly 1/8 the price of Opus-5.5, Pareto-optimal on cost, though its output token efficiency trails Astra at every effort tier.
- On maxbittker's Runescape Bench (tongue-in-cheek), GPT-6.1 Sol ranked second at about 10% of Astra's price.
- cedricchee tested 3D generation: quality close to Astra, roughly 30% faster, with an AA Index 4 points above GPT-6 Sol.
- Developer codestantine, on a self-built low-resource language translation benchmark, claims GPT-6.1 Sol beats Astra, with a bigger jump than going from 5.6 Sol to 6 Astra (shared by nickbaumann).
- Reddit user therealjerseytom said it "lives up to the hype," with token costs far below 6 Astra for the same effort—especially notable given 6 Astra's high cost and painful rework on failures.
- User chogku says the GPT-6.1 Sol + Opus 5.5 combo is high quality, cheap, and near-Astra in capability; rudrank shared it calling it "a dream come true."
Unconfirmed
- Independent evaluator BridgeBench measured GPT 6.1 Sol at only 3 points above GPT 6 Sol and 82 points behind GPT 6 Astra, a clear divergence from OpenAI's "near-Astra level intelligence" framing and other reviewers' positive results—so the model's true positioning remains contested.
- Huryn's additional n=3 reruns were still in progress, so final scores per tier may be updated.
Why it matters
- The shift in cost structure is the core takeaway of this round: multiple independent tests consistently show near-flagship performance delivered at 1/5 to 1/10 the price, which could reshape developers' default model choices and the economics of large-scale agent deployment.
- Third-party benchmarks don't fully agree (e.g., BridgeBench), a reminder that model capability should be assessed task by task rather than generalized.
- GPT-6.1 Sol near-Astra quality in 3D work while running ~30% faster — cedric_chee · 2026-09-30
- GPT-6.1 Sol Nears Astra 3D Quality at ~30% Faster, Gains 4 Points on AA Index — cedric_chee · 2026-09-30
- GPT-6.1-Sol Takes #2 on eyebench-v3, ~8x Cheaper Than Opus-5.5 — adonis_singh · 2026-09-30
- GPT-6.1-Sol Takes #2 on eyebench-v3, ~8x Cheaper Than Opus-5.5 — adonis_singh · 2026-09-30
- Cost Pareto-dominant, but second only to Astra on output-token efficiency — adonis_singh · 2026-09-30
- GPT 6.1 Sol Only +3 on BridgeBench, 82 Points Behind Astra: Benchmaxing Suspected — RexDouglass · 2026-09-30
- GPT-6.1 Sol scores 44 on real repos for just $6.56, near-top accuracy at 1/5 the cost — PawelHuryn · 2026-09-30
- GPT-6.1 Sol fixes 44 of 105 planted bugs for $6.56, matching Astra at a fraction of the cost — PawelHuryn · 2026-09-30
- GPT-6.1 Sol lands #2 on 'Runescape Bench' at ~10% the price of Astra — banteg · 2026-09-30
- GPT-6.1 Sol nearly matches flagship on 105 planted bugs, at ~1/5 the cost — PawelHuryn · 2026-09-30
- Bug Hunt Bench updates all effort tiers; GPT-6.1 Sol xhigh nearly free on 105 planted bugs — PawelHuryn · 2026-09-30
- Early GPT-6.1 Sol user report: similar work done at dramatically lower token cost than 6 Astra — therealjerseytom · 2026-09-30
- GPT-6.1 Sol + Opus 5.5 combo called near-Astra quality at a fraction of the cost — rudrank · 2026-10-01
- GPT-6.1 Sol beats Astra on low-resource translation, dev benchmark shows — nickbaumann_ · 2026-10-01
- GPT-6.1 Sol benchmarks across all effort levels land, making GPT-6 Astra hard to justify — PawelHuryn · 2026-10-01
- GPT-6.1 Sol Benchmarked Across All Effort Levels; Atra Faster but Pricier — PawelHuryn · 2026-10-01
Episode 2 · Claim of GPT-6.1 Sol Crushing Opus 5.5 Sparks Benchmark Skepticism (2026-09-30, 2 posts)
A Reddit post claimed OpenAI's internal benchmarks show GPT-6.1 Sol far ahead of Claude Opus 5.5, but a follow-up post mocked the claim as evidence of blind benchmark worship.
- OpenAI's internal benchmarks reportedly show GPT-6.1 Sol crushing Opus 5.5 — wilyi · 2026-09-30
- Fake 'GPT-6.1 Sol crushes Opus 5.5' meme mocks eval-driven hype — bigblueboo · 2026-10-01