105 planted bugs tested: GPT-6 Astra tops at 45, GPT-6 Sol shows big degradation

JohnMcKeownn · x · 2026-09-24

Pawel Huryn ran frontier coding models against 105 hidden bugs in two real repos: GPT-6 Astra (max) fixed 45, GPT-5.6 Sol 43.5, Opus 5.5 41.7, Muse Spark 1.3 32.2, and GPT-6 Sol only 29.3—a large degradation. API-equivalent cost varies widely too: GPT-6 Astra $33, Opus 5.5 $58.5, GPT-5.6 Sol over $95. The bug list stays private to keep the benchmark valid, but test coverage is published.

Related event: GPT-6 Astra Tops Bug Hunt Bench While Sol Version Regresses(2 posts)→

Original post →

More from coding & agent

coding & agent channel →