DeepSWE 1.1 Benchmark Results Announced

gabrielchua · x · 2026-07-10

The official results for the latest models on the DeepSWE 1.1 benchmark have been released, highlighting that the 5.6 Sol version is at the forefront when balancing performance against cost, output tokens, and agent steps.

Officials also noted that the 5.6 Terra and Luna versions perform exceptionally well, encouraging users to test them in real-world applications and provide feedback rather than focusing solely on benchmark scores.

Related event: GPT-5.6 Tops DeepSWE Leaderboard with Superior Cost-Efficiency(11 posts)→

Original post →

More from Models

Models channel →