Kimi K3 Matches Fable 5 in SWE Benchmarks at a Third of the Cost

Recent DeepSWE benchmark data reveals that the open-source Kimi K3 model performs on par with the closed-source Claude Fable 5 in software engineering tasks. The per-task correlation between the two is as high as 0.72, a record for models from different creators. Author @FinanceYF5 notes that cutting-edge open-source models are no longer six months behind proprietary ones, signaling a shift in the industry landscape.

Key Details and Performance Comparison

Regarding core pass metrics, Kimi K3 and Fable 5 differ by only 1 point on pass@1. However, Kimi K3 shows an advantage with larger sampling budgets, achieving 82.0% on pass@2 and 89.4% on pass@4, surpassing GPT-5.6 Sol. For programming languages, Fable 5 leads in Python, JavaScript, TypeScript, and Rust, while Kimi K3 outperforms in Go (79 vs. 71). Furthermore, their failure modes are nearly identical, with about 65% of failures classified as "near misses." They also maintain baselines well, with regression rates of 11% and 10% respectively. Since no extreme polarization was observed where one model consistently passes and the other fails, the analysis suggests the benchmark may be nearing saturation.

Compute Economics and Cost Advantages

Alongside comparable performance, Kimi K3 offers significant economic advantages. Data shows Kimi K3's single-run cost is only $4.65, compared to $13.41 for Fable 5. Calculated per $100 invested, Kimi K3 can solve 14.7 tasks, which is 2.8 times the throughput of Fable 5.

2026-07-21 ~ 2026-07-22 · 7 related posts

Full story(20 episodes)→

Primary sources