GPT-6 Sol Underperforms GPT-5.6 Max on DeepSWE, 68.8% vs 72.7%
banaxi-tech · reddit · 2026-09-23
A Reddit user compared GPT-6 Sol against the previous generation on the DeepSWE benchmark: GPT-5.6 Max scored 72.7% while GPT-6 Sol managed only 68.8%, a near 4-point regression. This matches widespread user complaints that GPT-6 Sol underperforms on non-coding tasks and barely engages in extended reasoning, fueling debate over what the new generation traded away.
More from Models
- Jev clone wave: dozens of open-source alternatives appear within a week of launch — airesearch12 · 2026-09-23
- DiffusionGemma-Jev lands in vLLM: single-step structured answers with confidence scores — bodonoghue85 · 2026-09-23
- Muse Spark 1.3 hands-on: great price, still can't write proofs — tianyin_xu · 2026-09-23
- $1000 bounty draws no takers: critics of French AI-novel detector can't cite one false positive — birchlse · 2026-09-23
- Hugging Face engineer: we need a decision model that auto-scales reasoning effort — mishig25 · 2026-09-23
- OpenAI reportedly plans Free, Prototype, and Accelerate tiers for its app-building platform — testingcatalog · 2026-09-23