Luna Max outperforms Sonnet in DeepSWE at 2.3% of the cost
reach_vb · x · 2026-08-17
On the DeepSWE v1.1 benchmark, GPT-5.6 Luna Max scores 13.3 points higher than Sonnet 5 Max while costing 44x less. DeepSWE evaluates coding agents on 113 original, long-horizon engineering tasks. Luna achieves a 67.2% pass rate at $0.61/task, performing 1.7 points lower than Gemini 3.7 Flash Medium (70% pass rate) but at 70% lower cost.
Related event: GPT-5.6 Luna Max Tops Sonnet 5 Max on DeepSWE at 1/44 the Cost(3 posts)→
More from Models
- DeepSeek adds peak/off-peak pricing: V4 Pro output $3.96 at peak, 2x old price off-peak — AccBalanced · 2026-08-17
- Anthropic's Internal "Model 2" Leaked, Outperforms Mythos 5 — koltregaskes · 2026-08-17
- Why run multiple models: The whole is greater than the sum of its parts — cantrell · 2026-08-17
- Dev take: Kimi K3 self-corrects too much, GLM 5.2 hits cognitive limits — tokumin · 2026-08-17
- Polymarket: OpenAI's 'Astra' model has 52% chance of release within a month — Polymarket · 2026-08-17
- Gemini 3.7 Flash Still Far Behind 3.1 Pro in Three.js Task — Able-Line2683 · 2026-08-17