Bug Hunt Benchmark: GPT-6.1 Sol matches or beats GPT-5.6 Sol at over 10x lower cost
PawelHuryn · x · 2026-10-07
Paweł Huryn tested Opus 5.5, Sonnet 5.5, and GPT-6.1 Sol on his Bug Hunt Benchmark: 105 bugs that frontier models missed in early 2026, across 2 real repos. Key findings:
- Opus 5.5 is significantly stronger than Opus 5 and nearly matches Fable 5.1 (41.7 vs 43) at two-thirds the cost ($58.53 vs $87.18). He recommends dropping Fable 5.1 unless budget is unlimited.
- Sonnet 5.5 costs half of Opus 5.5 on tokens ($2/$10 vs $4/$20 per MTok), but long agentic sessions are dominated by cache reads, priced identically at $0.20/MTok. Sonnet 5.5 (max) topped the benchmark at 51.3/105 after 1,497 iterations.
- GPT-6.1 Sol is slightly stronger than GPT-5.6 Sol yet over 10x cheaper on complex tasks: half-price tokens, 4x cheaper cache reads (framed as a 95% discount, possibly temporary). Much of the gap comes from fewer turns: 485 turns for GPT-5.6 xhigh vs 200 for GPT-6.1 on the same task.
- On launch day GPT-6.1 Sol ran nearly 2x slower than Opus 5.5; OpenAI's Tibo says it should get 2x faster within hours.
All data and benchmark measurements are free to read.
Related event: Independent Bug Hunt Benchmark Ranks Latest AI Models(3 posts)→
More from coding & agent
- Van Gogh Starry Night town generator adds real-time brushstroke flow fields — simonxxoo · 2026-10-07
- VC's experiment: ChatGPT and Claude flip-flop on retirement investing, showing why vertical AI wins — MartinGTobias · 2026-10-07
- AI reverse-engineers an open-source implementation of the entire Adobe Suite — wen_ragnarok · 2026-10-07
- Shipping an LLM Feature to the Public: 7 Guards That Weren't the Prompt — clementds · 2026-10-07
- Teknium fixes Hermes Agent bug that silently dropped lessons for user-owned skills — Teknium · 2026-10-07
- Java Vector API: Writing SIMD Directly Since JDK 16 to Unlock Single-Core Performance — lemire · 2026-10-07