Opus 5.5 nearly matches Fable 5.1 at two-thirds the cost in real-repo bug benchmark
PawelHuryn · x · 2026-10-07
Distrusting public leaderboards, Paweł Huryn tested Opus 5.5, Sonnet 5.5 and GPT-6.1 Sol on his Bug Hunt Benchmark (2 real repos, 105 bugs frontier models missed in early 2026):
- Opus 5.5 is far stronger than Opus 5 and nearly matches Fable 5.1 (41.7 vs 43) at two-thirds the cost ($58.53 vs $87.18) — he recommends dropping Fable 5.1 unless budget is unlimited
- Sonnet 5.5 costs half of Opus 5.5 on input/output ($2/$10 vs $4/$20 per MTok), but long agentic sessions are dominated by cache reads, priced identically at $0.20/MTok, erasing the gap
- Sonnet 5.5 (max) surprisingly won (51.3/105) by running 1,497 turns
- GPT-6.1 Sol is slightly stronger yet 10x+ cheaper than GPT-5.6 Sol on complex tasks: half-price tokens, 4x cheaper cache reads (a 95% discount that may be temporary); the cost gap mostly reflects fewer turns (200 vs 485)
- At launch GPT-6.1 Sol was 2x slower than Opus 5.5; OpenAI's Tibo says it should nearly double in speed within hours
All data and measurements are freely available.
Related event: Independent Bug Hunt Benchmark Ranks Latest AI Models(3 posts)→
More from coding & agent
- Shipping an LLM Feature to the Public: 7 Guards That Weren't the Prompt — clementds · 2026-10-07
- Teknium fixes Hermes Agent bug that silently dropped lessons for user-owned skills — Teknium · 2026-10-07
- Java Vector API: Writing SIMD Directly Since JDK 16 to Unlock Single-Core Performance — lemire · 2026-10-07
- An AI Agent Audits Its Own Memory File: 71 of 147 Rules Cited by Nothing — Most-Agent-7566 · 2026-10-07
- Has Anyone Actually Used a Personal AI Agent for the Full Job-Search Loop? — haseeb_heaven · 2026-10-07
- Veteran Dev: The Real Line Is Handing Your Entire Codebase to the Agent — erwinalp5 · 2026-10-07