Luna Outperforms Sonnet on DeepSWE at 44x Lower Cost
reach_vb · x · 2026-08-17
Data shows GPT-5.6 Luna Max scores 13.3 points higher than Sonnet 5 Max on the DeepSWE v1.1 benchmark while costing 44 times less. DeepSWE evaluates coding agents on 113 original, long-horizon engineering tasks. Luna achieves a score of 67.2% at $0.61 per task, outperforming Gemini 3.7 Flash Medium by 1.7 points at 70% lower cost.
Related event: GPT-5.6 Luna Max Tops Sonnet 5 Max on DeepSWE at 1/44 the Cost(3 posts)→
More from Models
- Ivo AI releases Ivo Sage, an open-source frontier legal model trained on River API — xiaosun86 · 2026-10-03
- AI slide design is even more conspicuous than AI writing — siglesias · 2026-10-03
- Humanlike Qwen3.8-27B LoRA 2.0 Adds Tool Calls, Fooled a Blind Judge 23.5% of the Time — kvyb · 2026-10-03
- Local AI community worries growing dependence on Claude and GPT strengthens closed labs — takoulseum · 2026-10-02
- Unreleased Gemini 4 Argon reportedly matches Claude's best on 3D game generation — 141_1337 · 2026-10-02
- Team ditches GPT-5.4 for GLM in production, sees faster, cheaper, more reliable results — ivan_bezdomny · 2026-10-02