Luna Outperforms Sonnet on DeepSWE at 44x Lower Cost

reach_vb · x · 2026-08-17

Data shows GPT-5.6 Luna Max scores 13.3 points higher than Sonnet 5 Max on the DeepSWE v1.1 benchmark while costing 44 times less. DeepSWE evaluates coding agents on 113 original, long-horizon engineering tasks. Luna achieves a score of 67.2% at $0.61 per task, outperforming Gemini 3.7 Flash Medium by 1.7 points at 70% lower cost.

Related event: GPT-5.6 Luna Max Tops Sonnet 5 Max on DeepSWE at 1/44 the Cost(3 posts)→

Original post →

More from Models

Models channel →