GPT-5.6 Luna MAX Shows Anomalous Results on DeepSWE

vikdean · reddit · 2026-07-11

Someone on Reddit mentioned the performance of GPT-5.6 Luna MAX - DeepSWE, stating that it is "very strong" under the max setting.

However, the author also pointed out that the numbers for the low and medium settings are quite strange. Such a massive gap between different effort levels has never been seen before on this benchmark, suggesting that the model's performance across varying reasoning intensities might be unstable or anomalously different.

Original post →

More from Models

Models channel →