Claude Sonnet 5.5 Launch: Near-Opus Intelligence but Token Efficiency Under Fire
Anthropic has released the mid-tier Claude Sonnet 5.5 (priced at half of Opus 5.5), and third-party benchmarks and hands-on feedback rolled in fast: intelligence close to the flagship Opus 5.5, but token consumption and cost efficiency emerged as the biggest controversies. The consensus: "smart but expensive" — whether to upgrade depends on your use case.
Confirmed
- Artificial Analysis intelligence index: 56 in max mode, just 2 points below Opus 5.5 (max); across five reasoning levels (max 56, xhigh 52, high 47, medium 41, etc.) pricing varies 18x, from $0.41 to $7.60 per task
- Terminal-Bench 4.0 score of 64%, a 50-point jump over Sonnet 5 (max)
- AA measured max-effort tasks averaging 193K output tokens — the highest of any tested model, roughly 7x GPT-6 Astra; Reddit user @Roflxd88 noted it produced 410 million cumulative tokens on the Intelligence Index, making it the "most verbose" model
- The @every team's hands-on testing: 30% faster than Sonnet 5, with costs down up to 30% in most workflows; Dan Shipper's private writing benchmark Dan's Editorial Checks (7 models, 5 tasks) scored it 65% overall with 90% on paragraph revision — writing even better than Opus 5.5
- @Duex's custom blind test (a Rust DEFLATE/zlib decompressor task): accuracy on par with Opus 5.5 at about a quarter of the price
- Reviewer Matthew Berman's early hands-on: essentially at Opus 5.5 level, 50% cheaper and faster
Unconfirmed
- @haider1 (an AI practitioner with 70K followers) tested it and questioned whether Sonnet 5.5 is "benchmark-chasing": impressive scores but noticeably more expensive and token-hungry, with cost and token efficiency on the AA index even worse than Sonnet 5
- Reddit user @Gohab2001 questioned why Anthropic released a model positioned this way (possibly aimed at specific scenarios)
Why it matters
- Claude models take four of the top five spots on the intelligence index (AA data cited by @Hesamation), underscoring Anthropic's strong position in model capability
- Sonnet 5.5 illustrates the "intelligence for tokens" trade-off: capability approaching the flagship while efficiency metrics regress versus the previous generation, adding a new case study to the cost-pricing debate around high-reasoning-effort modes
- The Every team's advice splits by scenario: worth upgrading for short-iteration feedback loops and planning-heavy collaboration; for coding, Kieran Klaassen thinks Sonnet 5 is still sufficient; Katie Parrott calls it a better collaborator but advises users to stay in control and know when to escalate
2026-09-29 ~ 2026-09-29 · 27 related posts
- Episode 1: Insiders claim OpenAI and Anthropic hold models far ahead of public releases(2026-09-24, 2 posts)
- Episode 2: Anthropic Reportedly Set to Ship Sonnet/Haiku 5.5 Ahead of OpenAI DevDay(2026-09-26, 2 posts)
- Episode 3: Claude Sonnet 5.5 leak: gradual rollout underway, launch imminent(2026-09-27, 20 posts)
- Episode 4: Insider Claims OpenAI Is Sitting on a Stronger Model(2026-09-28, 2 posts)
- Episode 5: Rumor: Anthropic ships Opus 5.5 and Sonnet 5.5 right before OpenAI DevDay(2026-09-29, 7 posts)
- Episode 6: Anthropic Launches Claude Sonnet 5.5: 30% Faster, Up to 30% Cheaper(2026-09-29, 58 posts)
- Episode 7: Claude Sonnet 5.5 Launch: Near-Opus Intelligence but Token Efficiency Under Fire(2026-09-29, 27 posts)
- Episode 8: Rumored benchmarks show Claude Sonnet 5.5 crushing GPT-6 sol(2026-09-29, 5 posts)
- Episode 9: Data comparison sparks debate: Sonnet 5.5 called a flop, GPT-6 Sol wins on value(2026-09-29, 2 posts)
Primary sources
- Sonnet 5.5 scores 56 on AA Index, 2 points off Opus 5.5—but with record 193k tokens per task — ArtificialAnlys ·
- Sonnet 5.5 lands at half Opus 5.5's price, Every team splits after a week of testing — every ·
- Sonnet 5.5 is a token-hungry monster: 193k output tokens, 2x GPT-6's cost per task — haider1 ·
- Early Tests: Sonnet 5.5 Matches Opus 5.5 at Half the Price and Much Faster — EricBuess · 2026-09-29
- [source] Sonnet 5.5 scores 56 on AA Index, 2 points off Opus 5.5—but with record 193k tokens per task — ArtificialAnlys · 2026-09-29
- Sonnet 5.5 jumps 50 points to 64% on Terminal-Bench 4.0, topping Opus 5.5 — ArtificialAnlys · 2026-09-29
- Claude Sonnet 5.5 Burns ~193k Output Tokens Per Task, 7x More Than GPT-6 Astra — ArtificialAnlys · 2026-09-29
- Artificial Analysis Publishes Full Eval Breakdown for Claude Sonnet 5.5 Across Reasoning Efforts — ArtificialAnlys · 2026-09-29
- Claude Sonnet 5.5 Efforts Span 18x Price Range: $0.41 to $7.60 Per Task on AA Index — ArtificialAnlys · 2026-09-29
- Four of the five smartest models on the AA Intelligence Index are now Claude — Hesamation · 2026-09-29
- Claude takes 4 of top 5 spots on AA Intelligence Index, Sonnet 5.5 beats Fable 5.1 — xiaosun86 · 2026-09-29
- Sonnet 5.5 sets Artificial Analysis record for most output tokens at lower cost — Gohab2001 · 2026-09-29
- Anthropic launches Claude Sonnet 5.5: 56 on AA Index, just 2 points behind Opus 5.5 — jarrodwatts · 2026-09-29
- [source] Sonnet 5.5 is a token-hungry monster: 193k output tokens, 2x GPT-6's cost per task — haider1 · 2026-09-29
- Blind test: Sonnet 5.5 matches Opus 5.5 on a from-scratch Rust decompressor at 1/4 the price — _Duex · 2026-09-29
- AiBreakfast: Sonnet 5.5 now beats Opus 5.5 and Fable 5.1 overall — AiBreakfast · 2026-09-29
- Sonnet 5.5 Ships With Opus-Level Writing Gains, But Mid-Tier Models May Be Dying — every · 2026-09-29
- Every's vibe check: Claude Sonnet 5.5 runs 30%+ faster, up to 30% cheaper than Sonnet 5 — danshipper · 2026-09-29
- Every team tests Sonnet 5.5: half Opus's token price, matches it on writing readability — kieranklaassen · 2026-09-29
- Vibe Check: Sonnet 5.5 Is a More Capable Partner Than Sonnet 5, If You Keep a Hand on the Wheel — kieranklaassen · 2026-09-29
- Sonnet 5.5 benchmarks near Opus 5.5 unevenly, while OpenAI keeps token-efficiency edge — cedric_chee · 2026-09-29
- Every's vibe check: Sonnet 5.5 is 30% faster and up to 30% cheaper than Sonnet 5 — every · 2026-09-29
- Every's Vibe Check: Sonnet 5.5 Fixes Claude Code Bug, 30% Faster With 30% Less Usage — every · 2026-09-29
- Sonnet 5.5 Tops Every's Readability Scores, Feels Like Opus 5.5 and Fast — every · 2026-09-29
- Early user impressions: Sonnet 5.5 faster and ~30% cheaper than 5.0, near Opus 5.5 — shaunralston · 2026-09-29
- Early hands-on says Sonnet 5.5 looks benchmaxxed: pricier and more token-hungry than Sonnet 5 — haider1 · 2026-09-29
- [source] Sonnet 5.5 lands at half Opus 5.5's price, Every team splits after a week of testing — every · 2026-09-29
- After a Week of Testing, Sonnet 5.5 Wins at Collaboration but Sonnet 5 Still Matches It on Coding — every · 2026-09-29
- Sonnet 5.5 logs 410M output tokens on Intelligence Index, the most verbose model yet — Roflxd88 · 2026-09-29
- Hands-on with Claude Sonnet 5.5: half the cost, 30% faster, the go-to execution model — 卡尔的AI沃茨 · 2026-09-29