SWE-Bench Pro V2 Hard nearly saturated at 98% on day one
141_1337 · reddit · 2026-09-23
A Reddit user highlighted that SWE-Bench Pro V2 Hard was already nearly saturated at 98% by top models on its very first day, meaning the benchmark can no longer meaningfully differentiate frontier coding models and underscoring the accelerating pace of benchmark saturation.
Related event: SWE-Bench Pro V2 Hard Hits 98% on Day One, Instantly Saturated(2 posts)→
More from Models
- Andriy Burkov: Codex Is Infinitely Faster Than Any Open-Weight Agent, But Priced Out of Reach — burkov · 2026-09-23
- Viral demo claims 'GPT-6' can drive browser Paint to draw, unverified — alexcovo_eth · 2026-09-23
- Opus 5.5-generated three.js spell demo wows with procedural VFX and sound — majidmanzarpour · 2026-09-23
- Early hands-on with rumored Opus 5.5 in Scenario's Blender plugin stuns users — repligate · 2026-09-23
- GPT-6 Sol's price cut means it should be compared to Sonnet, not Opus — TraditionalHome8852 · 2026-09-23
- User burns two OpenAI 20x accounts down to 13%, sizing up GPT-6 efficiency — jdjohnson · 2026-09-23