FULL STORY

GPT-6.1 Sol: Strong Independent Tests, Contested Internal Benchmarks

After GPT-6.1 Sol's launch, independent testers found performance near GPT-6 Astra at far lower cost. OpenAI's internal benchmarks claiming a big lead over Claude Opus 5.5 soon drew community skepticism.

2026-09-30 ~ 2026-10-01 · 2 episodes · 18 posts

Episode 1 · GPT-6.1 Sol benchmarks land: near-Astra performance at a fraction of the cost (2026-09-30, 16 posts)

Since GPT-6.1 Sol launched, multiple independent reviewers have tested it on coding, 3D generation, translation and other tasks. The mainstream verdict: performance close to GPT-6 Astra with dramatically lower token costs, giving it a clear edge in value for money—though a few benchmark results diverge from the official "near-Astra level" claim.

Confirmed

  • Pawel Huryn ran a blind test (Bug Hunt Bench) planting 105 bugs across 2 real code repositories: at max effort, GPT-6 Astra fixed 45 for $33, while GPT-6.1 Sol fixed 44 for just $6.56—roughly 1/5 of Astra's cost or less. He judged this one "the real deal," unlike GPT-6 Sol, which was seen as a watered-down GPT-5.6 Terra.
  • Huryn later published scores for all effort tiers (n=3 for max/xhigh, n=2 for the rest, with follow-up reruns ongoing), concluding that Astra is faster but pricier—given the costs, he sees no reason to keep using GPT-6 Astra.
  • adonissingh's eyebench-v3 leaderboard ranks GPT-6.1-Sol second, displacing Opus-5.5; it's about 3.8x cheaper than astra and roughly 1/8 the price of Opus-5.5, Pareto-optimal on cost, though its output token efficiency trails Astra at every effort tier.
  • On maxbittker's Runescape Bench (tongue-in-cheek), GPT-6.1 Sol ranked second at about 10% of Astra's price.
  • cedricchee tested 3D generation: quality close to Astra, roughly 30% faster, with an AA Index 4 points above GPT-6 Sol.
  • Developer codestantine, on a self-built low-resource language translation benchmark, claims GPT-6.1 Sol beats Astra, with a bigger jump than going from 5.6 Sol to 6 Astra (shared by nickbaumann).
  • Reddit user therealjerseytom said it "lives up to the hype," with token costs far below 6 Astra for the same effort—especially notable given 6 Astra's high cost and painful rework on failures.
  • User chogku says the GPT-6.1 Sol + Opus 5.5 combo is high quality, cheap, and near-Astra in capability; rudrank shared it calling it "a dream come true."

Unconfirmed

  • Independent evaluator BridgeBench measured GPT 6.1 Sol at only 3 points above GPT 6 Sol and 82 points behind GPT 6 Astra, a clear divergence from OpenAI's "near-Astra level intelligence" framing and other reviewers' positive results—so the model's true positioning remains contested.
  • Huryn's additional n=3 reruns were still in progress, so final scores per tier may be updated.

Why it matters

  • The shift in cost structure is the core takeaway of this round: multiple independent tests consistently show near-flagship performance delivered at 1/5 to 1/10 the price, which could reshape developers' default model choices and the economics of large-scale agent deployment.
  • Third-party benchmarks don't fully agree (e.g., BridgeBench), a reminder that model capability should be assessed task by task rather than generalized.

Episode 2 · Claim of GPT-6.1 Sol Crushing Opus 5.5 Sparks Benchmark Skepticism (2026-09-30, 2 posts)

A Reddit post claimed OpenAI's internal benchmarks show GPT-6.1 Sol far ahead of Claude Opus 5.5, but a follow-up post mocked the claim as evidence of blind benchmark worship.