GPT-6 Sol Scales Near-Linearly With Effort but Lags in Bug Hunt

Developer tests show GPT-6 Sol scores scale almost linearly with effort levels. However, Bug Hunt Bench's blind evaluation of 105 real bugs finds GPT-6 Sol (max) still trails Opus 5.5 (medium).

2026-09-23 ~ 2026-09-23 · 2 related posts

Full story(4 episodes)→