105 planted bugs tested: GPT-6 Astra tops at 45, GPT-6 Sol shows big degradation
JohnMcKeownn · x · 2026-09-24
Pawel Huryn ran frontier coding models against 105 hidden bugs in two real repos: GPT-6 Astra (max) fixed 45, GPT-5.6 Sol 43.5, Opus 5.5 41.7, Muse Spark 1.3 32.2, and GPT-6 Sol only 29.3—a large degradation. API-equivalent cost varies widely too: GPT-6 Astra $33, Opus 5.5 $58.5, GPT-5.6 Sol over $95. The bug list stays private to keep the benchmark valid, but test coverage is published.
Related event: GPT-6 Astra Tops Bug Hunt Bench While Sol Version Regresses(2 posts)→
More from coding & agent
- One Stripe Engineer Merged 600 AI-Written Changes in Six Months, One Rollback — victor_explore · 2026-09-26
- OpenClaw Creator steipete on How a WhatsApp Bot Became a Top Open-Source Agent Project — steipete · 2026-09-26
- Agent Tincan: Open-Source Tool Lets Your AI Agents Talk to Each Other Over Tailscale — Scobleizer · 2026-09-26
- 5 GPUs power a real-time AI panel where multiple avatars debate and take live questions — Opening-Trip9912 · 2026-09-26
- Two Prompts, One Showreel: Claude Code Makes a Full Motion Graphics Video — tristanbob · 2026-09-26
- Open-source hallpass adds live per-user authorization checks to MCP server write operations — Adorable-Algae6903 · 2026-09-26