GPT-6 Sol Underperforms GPT-5.6 Max on DeepSWE, 68.8% vs 72.7%

banaxi-tech · reddit · 2026-09-23

A Reddit user compared GPT-6 Sol against the previous generation on the DeepSWE benchmark: GPT-5.6 Max scored 72.7% while GPT-6 Sol managed only 68.8%, a near 4-point regression. This matches widespread user complaints that GPT-6 Sol underperforms on non-coding tasks and barely engages in extended reasoning, fueling debate over what the new generation traded away.

Original post →

More from Models

Models channel →