DeepSeek-V3 Beats GPT-4o and Claude 3.5 Sonnet on DeepSWE Benchmark
zainhas · x · 2026-08-15
According to a user evaluation, DeepSeek-V3 (Pro 0813) achieves a higher pass@4 score on the DeepSWE benchmark compared to GPT-4o and Claude 3.5 Sonnet (referred to as Fable 5 in the post). The model is noted for having moments of brilliance.
More from Models
- DeepSeek-V4 Pro and Fable show lowest task correlation — zainhas · 2026-08-15
- DeepSeek-V4 Pro beats Sol and Fable in coding tasks — zainhas · 2026-08-15
- 27B Qwen 3.8 Beats Rumored 1-5T Param Opus 4.6 on All Benchmarks — zainhas · 2026-08-15
- 90% of Users Don't Need SOTA Models; Google Targets the Mass Market — haider1 · 2026-08-15
- Comparison of Veo, H3, and Seedance Video Generation — Then_Editor_7958 · 2026-08-15
- Opus model accused of rambling instead of pruning docs — JasonBotterill · 2026-08-15