DeepSeek-V3 Beats GPT-4o and Claude 3.5 Sonnet on DeepSWE Benchmark

zainhas · x · 2026-08-15

According to a user evaluation, DeepSeek-V3 (Pro 0813) achieves a higher pass@4 score on the DeepSWE benchmark compared to GPT-4o and Claude 3.5 Sonnet (referred to as Fable 5 in the post). The model is noted for having moments of brilliance.

Original post →

More from Models

Models channel →