Opus 5.5 sweeps GPT-6 Sol on every benchmark, but Sol runs tasks for $1.06
Vast-Grapefruits · reddit · 2026-09-23
A Redditor compiled a head-to-head of Anthropic's Opus 5.5 and OpenAI's GPT-6 Sol, which shipped the same day.
- Raw scores: Opus 5.5 wins every benchmark in the table.
- Surprise: On GDPval, the "real office work" eval, Sol scores roughly 100 Elo below the GPT-5.6 Sol it replaces — Artificial Analysis attributes the drop to weaker presentation and deliverables skipping required task parts, not reasoning failures.
- Cost edge: Sol costs half per token, and at max effort Opus 5.5 emits 119K output tokens per task vs 31K for Sol, which runs the full index for about $1.06 per task.
The author is waiting for real-world feedback over the coming days.
Related event: Opus 5.5 Sweeps Benchmarks Over GPT-6 Sol, Cost Remains the Debate(5 posts)→
More from Models
- Dev says Opus 5.5 is nowhere near Fable 5.1 for hard coding problems — bindureddy · 2026-09-23
- Opus 5.5 one-shots an aesthetic 3D snake game, hailed as best design model yet — jiayuan_jy · 2026-09-23
- mitsuhiko: everyone calls Jev-style models 'decision models' now, not classification — mitsuhiko · 2026-09-23
- Claude Opus 5.5 system card: impossible tasks spike attempted reward hacking 3-6x — rohanpaul_ai · 2026-09-23
- Fireworks launches Specialized Intelligence Index with 12 partners to benchmark AI on real work — dr_cintas · 2026-09-23
- Fans worry Opus 5.5 leapfrogs Astra 6 as pressure mounts on OpenAI — rickasaurus · 2026-09-23