GPT-6 Astra Tops Bug Hunt Bench While Sol Version Regresses
A real-world benchmark of 105 hidden bugs shows GPT-6 Astra leading with 45 fixes, while GPT-6 Sol regressed sharply, fixing only 29 bugs compared to 43.5 by its predecessor GPT-5.6 Sol.
2026-09-24 ~ 2026-09-25 · 2 related posts
- Episode 1: OpenAI Launches GPT-6 Sol and Luna at Half Price; Big Hallucination Gains but Not a Clean Sweep(2026-09-23, 92 posts)
- Episode 2: GPT-6 Sol/Luna launch splits opinion: half-price but dumber?(2026-09-23, 9 posts)
- Episode 3: LMArena Opens GPT-6 Sol and Luna for Public Testing(2026-09-23, 2 posts)
- Episode 4: Four New Models Reshape the Cost-Performance Frontier, Opus 5.5 Tops Intelligence Index(2026-09-23, 2 posts)
- Episode 5: OpenAI and Anthropic Slash Prices Within 90 Minutes, Kicking Off a New Model Price War(2026-09-23, 7 posts)
- Episode 6: Tests Show GPT-6 Sol Scales Nearly Linearly With Effort(2026-09-23, 3 posts)
- Episode 7: GPT-6 Astra Tops Bug Hunt Bench While Sol Version Regresses(2026-09-24, 2 posts)
- Episode 8: Leaks Map GPT-6 Codenames to GPT-5.6 Versions(2026-09-24, 3 posts)
- Episode 9: Bug Hunt Bench Ranks Frontier Coding Models on 105 Real Bugs with 200x Cost Gap(2026-09-24, 3 posts)
- Episode 10: OpenAI's efficient GPT-6 models seen as freeing compute for bigger frontier models(2026-09-25, 2 posts)
- 105 planted bugs tested: GPT-6 Astra tops at 45, GPT-6 Sol shows big degradation — JohnMcKeownn · 2026-09-24
- Third-party bench claims GPT-6 Sol fixed 29/105 bugs vs 43 for GPT-5.6 Sol — ssh4net · 2026-09-25