Bug Hunt Bench: GPT-6 Sol (max) matches GPT-5.6 medium but trails Opus 5.5
PawelHuryn · x · 2026-09-23
New data from Bug Hunt Bench (blind-graded, 105 planted real bugs):
- GPT-6 Sol (max) ≈ GPT-5.6 Sol (medium)
- GPT-6 Sol (max) < Opus 5.5 (medium)
The board also plots score vs cost (200x spread, log axis) and time (<7x spread). Run by Pawel Huryn of The Product Compass.
More from coding & agent
- Alibaba Launches Qwen Intelligence With 3 SOTA Mobile Agents, 90% End-to-End Success Rate — Alibaba_Qwen · 2026-09-23
- Deleting Three Agents for 40 Lines of Python Cut Latency 9s to 800ms and Cost 75% — Prestigious_Style267 · 2026-09-23
- Open-Source Digital Twins: Stateless Fake SaaS Services for Testing Gmail/Slack Agents — aofu_dev · 2026-09-23
- Automation vs Agentic AI: Most 'Agents' Are Just Scripts, LLM Only Fills the Gaps — forevergeeks · 2026-09-23
- Agentic DORA metrics: tracking PR start-to-merge to measure AI coding agents — vincent_koc · 2026-09-23
- Replacing Claude & Chrome with Strawberry, an AI browser with built-in agents — damienghader · 2026-09-23