Third-party bench claims GPT-6 Sol fixed 29/105 bugs vs 43 for GPT-5.6 Sol
ssh4net · x · 2026-09-25
- StatsWire reports that on Bug Hunt Bench, GPT-6 Sol fixed only 29 of 105 bugs, while the previous GPT-5.6 Sol fixed 43.
- Takeaway: OpenAI allegedly cut price and capability at the same time—cheaper, but clearly worse. (Third-party eval, unverified.)
Related event: GPT-6 Astra Tops Bug Hunt Bench While Sol Version Regresses(2 posts)→
More from Models
- Claude Opus 5.5 creates a 5-minute animated UMAP explainer from a phone prompt — goodside · 2026-09-25
- User has Claude write a history of economic thought while he slept, economists impressed — soumitrashukla9 · 2026-09-25
- Podcast: Opus 5.5 Beats Fable 5.1 on GDPval at 40% Lower Cost, Tops Code Arena — thursdai_pod · 2026-09-25
- Claude Opus 5.5 and GPT-6 Sol Shipped 101 Minutes Apart, Both Cheaper — thursdai_pod · 2026-09-25
- Embedding test: RAG ranks the no-refund policy first, showing models still matter — galratner · 2026-09-25
- Meme: doing your taxes with Mistral could cost you a $20k IRS visit — Rasmic · 2026-09-25