Bug Hunt Benchmark: Sonnet 5.5 (max) wins at 51.3/105 while GPT-6.1 Sol undercuts GPT-5.6 Sol by 10x
PawelHuryn · x · 2026-10-07
Paweł Huryn tested the newly released models on his Bug Hunt Benchmark — 2 real repos with 105 bugs frontier models missed in early 2026 — and calculated what each subscription plan really buys in API value.
Key findings:
- Opus 5.5 is significantly stronger than Opus 5 and nearly matches Fable 5.1 (41.7 vs 43) at two-thirds the cost ($58.53 vs $87.18); he recommends dropping Fable 5.1 unless budget is unlimited
- Sonnet 5.5 costs half of Opus 5.5 on input/output ($2/$10 vs $4/$20 per MTok), but long agentic sessions are dominated by cache reads, priced identically at $0.20/MTok; the max tier won with 51.3/105 over 1,497 runs
- GPT-6.1 Sol is slightly stronger on complex tasks yet 10x+ cheaper than GPT-5.6 Sol: half-price tokens, 4x cheaper cache reads (a 95% discount OpenAI says may be temporary); the gap largely comes from fewer turns (200 vs 485)
- On launch day GPT-6.1 Sol was 2x slower than Opus 5.5; OpenAI says speeds should roughly double shortly
All data and benchmark measurements are freely available.
Related event: Independent Bug Hunt Benchmark Ranks Latest AI Models(3 posts)→
More from coding & agent
- GPT-6-luna Unlocks More Reasoning Tokens via API: ~18k Tokens Scores ~80.5% on Terminal-Bench — LysandreJik · 2026-10-07
- Encrypted prompt injection: one Copilot model leaked secrets in half the tests — Haunting_Ganache_850 · 2026-10-07
- Enterprise AI rolls out backward: chatbots are the finish line, not the start — shashib · 2026-10-07
- Gradio's ML Intern can now build Gradio apps from a single prompt — Gradio · 2026-10-07
- Bug Hunt Benchmark retest: GPT-6.1 Sol recovers, Muse still cheapest strong agent — PawelHuryn · 2026-10-07
- One Claude skill plus Scenario MCP automates the entire video-editing workflow — smtabatabaie · 2026-10-07