GPT-6 sol scores worse than GPT-5.6 sol on DeepSWE, sparking benchmark trust concerns
adonis_singh · x · 2026-09-23
- weswinder reports GPT-6 sol performs worse than GPT-5.6 sol on DeepSWE, the software-engineering agent benchmark, despite being cheaper.
- He finds it odd that a new model isn't better on every metric and suspects the vendor is trying to bury the benchmark.
- adonissingh amplifies: "do not trust it," suggesting DeepSWE scores may no longer reflect real capability.
More from Models
- Xiaomi launches open-source MiMo-V2.6 Pro & Flash, tops AI index at 46 among open models — burny_tech · 2026-09-23
- OfirPress: benchmark gap comes from a 166-task subset and pass-rate vs full-completion metrics — OfirPress · 2026-09-23
- Cline claims MiMo-v2.6-Pro is now the top open-weights model on AI index — burny_tech · 2026-09-23
- 'Sorry about that model': Opus 5.5 joke thread hails a huge jump over Opus 5 — repligate · 2026-09-23
- Claude Opus 5.5 wows users with creative output: 'draw everything you want' — repligate · 2026-09-23
- Ofir Press explains why third-party agent benchmark scores look higher: subset + different metric — OfirPress · 2026-09-23