Bug Hunt Benchmark: Opus 5.5 nears Fable 5.1 at 2/3 cost; GPT-6.1 Sol 10x cheaper
PawelHuryn · x · 2026-10-02
Paweł Huryn ran his Bug Hunt Benchmark (2 repos, 105 bugs frontier models missed in early 2026) to compute the real API value of AI subscription plans:
Model results
- Opus 5.5 is significantly stronger than Opus 5 and nearly matches Fable 5.1 (41.7 vs 43) at two-thirds the cost ($58.53 vs $87.18); he recommends dropping Fable 5.1 unless budget is unlimited
- Sonnet 5.5 costs half of Opus 5.5 on tokens ($2/$10 vs $4/$20 per MTok), but in long agentic sessions cache reads dominate and both are $0.20/MTok; Sonnet 5.5 (max) topped the benchmark at 51.3/105 in 1,497 turns
- GPT-6.1 Sol is slightly stronger than GPT-5.6 Sol yet over 10x cheaper on complex tasks (half input/output price, 4x cheaper cache reads — the 95% discount may be temporary); the gap mainly comes from turns (200 vs 485)
Subscription value
- Unlike OpenAI/Anthropic, xAI's small plans ($30 SuperGrok, $15 Muse) suffice for real agentic work; SuperGrok also includes the Grok Bot
- GPT-6.1 Sol launched nearly 2x slower than Opus 5.5; OpenAI's Tibo said it would get 2x faster within hours
Related event: Bug Hunt Benchmark Tests Real Value of AI Subscription Plans(2 posts)→
More from Models
- Cloudflare open-sources decision model Clef, outperforming Jev within weeks — michellechen · 2026-10-02
- GPT-6.1 Sol nearly matches Astra at a quarter of the price on RareBench — danielmckinn0n · 2026-10-02
- New results from PostTrainBench v1.2 are in — mariofilhoml · 2026-10-02
- Opus 5.5 on /high burns through Claude's weekly limit in two days — rudrank · 2026-10-02
- VLM Run launches SystemOne API: typed, calibrated answers in one forward pass — spillai · 2026-10-02
- Document classification bake-off: VLM is ~7x more token-efficient than OCR — spillai · 2026-10-02