Personal Benchmark: Opus 5.5 Hits Fable Level, GPT-5.6-Sol Still Leads

danshipper · x · 2026-09-23

Dan Shipper of Every points to @hammermt's personal benchmark Mike's Checks (9 models, 15 real work tasks) for a real-world read on Opus 5.5.

Highlights: Opus 5.5 scores 80% overall, matching Claude Fable 5 (79%) on many tasks, with 100% on pptx and roleplay-interview, though it timed out on a few tasks. GPT-5.6-Sol leads at 84%, Grok-4.6 hits 83%, and DeepSeek-V4-Flash trails at 58%.

The benchmark answers "which model is best at my job," scored via scripted checks plus an LLM judge, not a general capability claim.

Related event: Opus 5.5 Wins Praise for Speed and Efficiency in Same-Day Showdown With GPT-6 Sol(23 posts)→

Original post →

More from Models

Models channel →