FULL STORY

Claude Sonnet 5.5: Half-Price Launch to Benchmark Dominance

Anthropic launched Claude Sonnet 5.5 on Sept 29 at half price. Third-party leaks and benchmarks soon showed it beating GPT-6 sol and nearly matching Opus, though with top token usage.

2026-09-29 ~ 2026-09-29 · 3 episodes · 28 posts

Episode 1 · Anthropic Launches Claude Sonnet 5.5: Near-Flagship at Half the Token Price (2026-09-29, 11 posts)

On 09-29 Anthropic released Claude Sonnet 5.5, the second model in the Claude 5.5 family, officially described as a clear upgrade over Sonnet 5 with roughly 30% faster speed. Positioned as the "middle sibling" after Opus 5.5, multiple reviews conclude that its mid-tier pricing now approaches flagship capability.

Confirmed

  • Every founder Dan Shipper published a full vibe check review (covering coding, writing, design, and knowledge work): his conclusion is that Sonnet 5.5 is smarter and more efficient than Sonnet 5, about 30% faster, with task costs down as much as 30% in most work scenarios.
  • On Shipper's private writing benchmark based on real work, Dan's Editorial Checks (7 models, 5 tasks), Sonnet 5.5 scored 65% overall and 90% on paragraph revision—outperforming even Opus 5.5 in writing, and topping the readability rankings.
  • For coding, test cases show Sonnet 5.5 successfully fixed a bug in Claude Code; Shipper also said it retains Opus 5.5's "feel," with clear writing and fast speed.
  • Token pricing is half that of Opus 5.5, but Every's team notes that high-effort mode burns tokens quickly, so per-task costs aren't necessarily halved.

Not yet confirmed

  • Reviewer Matthew Berman said after early hands-on use that Sonnet 5.5 basically matches Opus 5.5 quality while being noticeably faster, with several demo videos (including a 3D simulation example)—though this is personal early experience.
  • User @shaunralston shared initial impressions that it is faster and smarter than Sonnet 5.0, close to Opus 5.5 capability and about 30% cheaper; currently a personal impression without official data to back it.

Why it matters

  • Multiple independent reviews converge on the same conclusion: a mid-tier model can now deliver near-flagship—or even locally superior—capability at half the token price and up to 30% lower task cost, a direct cost reduction for usage-billed developers and knowledge workers.
  • The overtake on writing tasks suggests the tiered positioning within the Claude 5.5 family is breaking down—"buy the flagship" is no longer necessarily the optimal choice for capability.

Episode 2 · Rumored benchmarks show Claude Sonnet 5.5 crushing GPT-6 sol (2026-09-29, 5 posts)

On September 29, multiple third-party sources simultaneously leaked claims that Anthropic's Claude Sonnet 5.5 significantly outperforms OpenAI's GPT-6 sol in benchmarks, sparking community debate about OpenAI's competitiveness. All figures currently come from unofficial channels and have not been confirmed by either party.

Confirmed

  • X user ns123abc (84K followers) posted screenshots claiming GPT-6 sol was soundly beaten by Claude Sonnet 5.5 in comparisons (their words: "brutally mogged").
  • LuminaBench leaked that Claude Sonnet 5.5 scored 56 on the Artificial Analysis Intelligence Index, versus 48 for the competing model Sol.
  • Another user claims Sonnet 5.5 scored 56 on the Intelligence Index—second only to Claude Opus 5.5 and above GPT-6 Astra.
  • The account AIScreening posted third-party test screenshots showing Sonnet 5.5 well ahead of GPT-6 sol and very close to GPT-6 Astra.
  • Another quoted post says Claude Sonnet 5.5 Max ranks second on the leaderboard, above GPT-6 Astra Max and just slightly below Opus-5.5; commenters see OpenAI as being in trouble.

Unconfirmed

  • All benchmark scores and rankings come from leaks and third-party reposts; test baselines, specific questions, and version details are unclear, with no official confirmation from Anthropic, OpenAI, or Artificial Analysis.
  • Full details of m1's evaluation methodology and source screenshots have not been disclosed.

Why it matters

  • If the leaks hold up, it means Anthropic's mid-tier Sonnet 5.5 could suppress OpenAI's GPT-6 sol and close in on GPT-6 Astra—directly affecting user choices and market narratives between the two vendors.
  • Multiple independent accounts spreading similar conclusions on the same day suggests that, even if the numbers are imprecise, community anxiety over the declining relative standing of OpenAI's frontier models is intensifying.

Episode 3 · Sonnet 5.5 Nearly Matches Opus 5.5 but Sets Token Consumption Record (2026-09-29, 12 posts)

Following Anthropic's release of Claude Sonnet 5.5, third-party evaluator Artificial Analysis published first benchmark data: the model scored 56 on the AA Intelligence Index, just 2 points below Opus 5.5 (max), with particularly strong performance under max effort; its Terminal-Bench 4.0 score of 64% is a 50-point improvement over Sonnet 5 (max). But while approaching flagship performance, its token consumption and cost are notably high — the core controversy of this evaluation round.

Confirmed

  • Intelligence Index score of 56, only 2 points behind Opus 5.5 (max); AA has published complete breakdowns of Intelligence Index results across reasoning effort levels.
  • At max effort, it averages about 193,000 output tokens per task — the highest of all tested models and roughly 7x GPT-6 Astra; compared with 119,000 for Opus 5.5 max, far exceeding sibling models.
  • Terminal-Bench 4.0 score of 64%, up 50 percentage points from Sonnet 5 (max).
  • AA's model comparison page shows the family's five reasoning effort tiers (max/xhigh/high/medium and below) scoring 56/52/47/41 points, with per-task costs ranging from $0.41 to $7.60 — an 18x spread.

Why it matters

  • Sonnet 5.5's mid-tier positioning near flagship performance reflects Anthropic's tiered reasoning-effort strategy, but the heavy consumption means real-world costs could double: @haider1's cost analysis calls it a "token guzzler," with per-task costs roughly 2x GPT-6.
  • For token-billed developers, Intelligence Index and unit cost must be weighed together, and AA's tiered breakdown offers a direct reference for model selection.