GPT-5.6 Luna MAX Shows Anomalous Results on DeepSWE
vikdean · reddit · 2026-07-11
Someone on Reddit mentioned the performance of GPT-5.6 Luna MAX - DeepSWE, stating that it is "very strong" under the max setting.
However, the author also pointed out that the numbers for the low and medium settings are quite strange. Such a massive gap between different effort levels has never been seen before on this benchmark, suggesting that the model's performance across varying reasoning intensities might be unstable or anomalously different.
More from Models
- Users say GPT-5.6 Ultra feels like extra token burn with little visible gain — CtrlAltDwayne · 2026-07-21
- LWiAI Podcast #252: OpenAI Launches GPT-5.6, LLM Pricing War Intensifies — Last Week in AI · 2026-07-21
- Early Gemini 3.6 Flash outputs look fast but weak on frontend and spatial reasoning — max_paperclips · 2026-07-21
- Anthropic removes Fable’s access deadline, but users say it was nerfed — oykun · 2026-07-21
- Kimi K3 retakes first place on DesignArena’s frontend web app benchmark — rohanpaul_ai · 2026-07-21
- Last Week in AI roundup covers Claude Sonnet 5, LongCat 2.0, and new agent benchmarks — Last Week in AI · 2026-07-21