Personal Benchmark: Opus 5.5 Hits Fable Level, GPT-5.6-Sol Still Leads
danshipper · x · 2026-09-23
Dan Shipper of Every points to @hammermt's personal benchmark Mike's Checks (9 models, 15 real work tasks) for a real-world read on Opus 5.5.
Highlights: Opus 5.5 scores 80% overall, matching Claude Fable 5 (79%) on many tasks, with 100% on pptx and roleplay-interview, though it timed out on a few tasks. GPT-5.6-Sol leads at 84%, Grok-4.6 hits 83%, and DeepSeek-V4-Flash trails at 58%.
The benchmark answers "which model is best at my job," scored via scripted checks plus an LLM judge, not a general capability claim.
More from Models
- Anthropic Launches Claude Opus 5.5: Fable 5.1-Level Performance at 40% Lower Cost — aigclink · 2026-09-23
- Heavy Opus 5.5 All-Day Use Barely Dents Usage Limits, Users Report — altryne · 2026-09-23
- Show today's LLMs to experts 10 years ago and they'd call it AGI — JacksonKernion · 2026-09-23
- Tired of running out of credits, this Reddit user says cheap model swarms work surprisingly well — AnotherWallace · 2026-09-23
- User complains Anthropic's new model auto-routes to fable, stays locked on cyber — BLUECOW009 · 2026-09-23
- Unconfirmed: Qwen4-35B-A3B reportedly being tested, not yet announced — AIFlow_ML · 2026-09-23