DeepSeek-V4 Pro beats Sol and Fable in coding tasks
zainhas · x · 2026-08-15
In a head-to-head benchmark on software engineering (DeepSWE) tasks, DeepSeek-V4 Pro achieved an 88.5% pass@4 accuracy, outperforming both Fable 5 and GPT 5.6 Sol. It also offers massive cost savings at $0.24 per task, making it 35x and 90x cheaper than Sol and Fable respectively.
More from Models
- Anthropic's next model won't ship externally; closed-source AI progress said to be paused — bindureddy · 2026-08-15
- Zhipu AI Releases GLM-5.3: 743B Base, Focus on Coding and Cyber Defense — max_paperclips · 2026-08-15
- DeepSeek-V4 Pro benchmark: Top performance at ultra-low cost — zainhas · 2026-08-15
- DeepSeek-V4 Pro and Fable show lowest task correlation — zainhas · 2026-08-15
- DeepSeek-V4 Pro coding analysis: Stable but weaker on Rust — zainhas · 2026-08-15
- Qwen3.8-Max launches on Together AI with 2.4T parameters — togethercompute · 2026-08-15