GPT-5.6 Outperforms Sonnet on DeepSWE at Lower Cost
reach_vb · x · 2026-08-16
Benchmarks indicate GPT-5.6 Luna Max scores 13.3 points higher than Sonnet 5 Max on DeepSWE v1.1 while being 44x cheaper. DeepSWE evaluates coding agents on 113 original, long-horizon engineering tasks.
More from Models
- User Test: Minimal Difference Between Google AI 3.6 and 3.7 Flash — gaganghotra_ · 2026-08-17
- Current models may lack strong covert capabilities, but doubts remain — TheZvi · 2026-08-17
- Pre-training defines AI's ceiling, and current limits are too small — pmddomingos · 2026-08-17
- Grok's looser safety offers better UX than Claude — billyjhowell · 2026-08-16
- Claude criticized for overusing paraprosdokians in marketing copy — churchkey · 2026-08-16
- Qwen3.8-27B Gets IQ4_XS Quantization for 16GB GPUs — Johnny_Rell · 2026-08-16