Luna Outperforms Sonnet on DeepSWE at 44x Lower Cost
reach_vb · x · 2026-08-17
Data shows GPT-5.6 Luna Max scores 13.3 points higher than Sonnet 5 Max on the DeepSWE v1.1 benchmark while costing 44 times less. DeepSWE evaluates coding agents on 113 original, long-horizon engineering tasks. Luna achieves a score of 67.2% at $0.61 per task, outperforming Gemini 3.7 Flash Medium by 1.7 points at 70% lower cost.
Related event: GPT-5.6 Luna Max Tops Sonnet 5 Max on DeepSWE at 1/44 the Cost(3 posts)→
More from Models
- Qwen3.8-27B hits 206 tok/s on single RTX 5090 via SGLang — StefanoGogioso · 2026-08-17
- antirez Optimizes DwarfStar: 170 t/s Generation and 22k tokens/s Prefill on Station — antirez · 2026-08-17
- OpenAI Introduces Tiered Access and Launches GPT-5.6-Cyber Security Model — dl_weekly · 2026-08-17
- Alibaba Cloud still offers cheap DeepSeek models — tobowers · 2026-08-17
- Claude Personification Moment: Rejecting Users and Judging Intentions — ctjlewis · 2026-08-17
- Qwen3.8 Benchmarks: MTP Settings Impact Throughput, Q4 Outperforms Q8 — New-Inspection7034 · 2026-08-17