Kimi K3 Matches Fable 5 in SWE Benchmarks at a Third of the Cost
Recent DeepSWE benchmark data reveals that the open-source Kimi K3 model performs on par with the closed-source Claude Fable 5 in software engineering tasks. The per-task correlation between the two is as high as 0.72, a record for models from different creators. Author @FinanceYF5 notes that cutting-edge open-source models are no longer six months behind proprietary ones, signaling a shift in the industry landscape.
Key Details and Performance Comparison
Regarding core pass metrics, Kimi K3 and Fable 5 differ by only 1 point on pass@1. However, Kimi K3 shows an advantage with larger sampling budgets, achieving 82.0% on pass@2 and 89.4% on pass@4, surpassing GPT-5.6 Sol. For programming languages, Fable 5 leads in Python, JavaScript, TypeScript, and Rust, while Kimi K3 outperforms in Go (79 vs. 71). Furthermore, their failure modes are nearly identical, with about 65% of failures classified as "near misses." They also maintain baselines well, with regression rates of 11% and 10% respectively. Since no extreme polarization was observed where one model consistently passes and the other fails, the analysis suggests the benchmark may be nearing saturation.
Compute Economics and Cost Advantages
Alongside comparable performance, Kimi K3 offers significant economic advantages. Data shows Kimi K3's single-run cost is only $4.65, compared to $13.41 for Fable 5. Calculated per $100 invested, Kimi K3 can solve 14.7 tasks, which is 2.8 times the throughput of Fable 5.
2026-07-21 ~ 2026-07-22 · 7 related posts
- Episode 1: Kimi K3 Tops Frontend Web App Arena with Enhanced English Skills(2026-07-21, 3 posts)
- Episode 2: Kimi K3 Jumps to 4th on Agent Arena Leaderboard(2026-07-21, 6 posts)
- Episode 3: Moonshot Releases 2.8 Trillion Parameter Open-Weight Model Kimi K3(2026-07-21, 8 posts)
- Episode 4: Kimi K3 Matches Fable 5 in SWE Benchmarks at a Third of the Cost(2026-07-21, 7 posts)
- Episode 5: Kimi K3 Sets New Open-Source ECI Record but Still Lags Behind(2026-07-22, 3 posts)
- Episode 6: Kimi K3 Accused of Gaming Benchmarks Instead of Solving Problems(2026-07-22, 2 posts)
- Episode 7: Kimi K3 Ranks Second on AA-Briefcase but with High Costs and Long Runtimes(2026-07-22, 6 posts)
- Episode 8: Kimi K3 Enters Top-Tier AI Model Ranks in Benchmark Tests(2026-07-22, 4 posts)
- Episode 9: Kimi K3 shifts attention from scale to architecture(2026-07-27, 25 posts)
- Episode 10: TokenSpeed Enables Kimi K3 Support on NVIDIA and AMD Platforms(2026-07-27, 2 posts)
- Episode 11: Moonshot's Kimi K3 Launches on Nebius with 1M Context(2026-07-27, 3 posts)
- Episode 12: SGLang Day-0 Support for Kimi K3 Boosts Throughput to 423 tok/s(2026-07-28, 8 posts)
- Episode 13: Moonshot's Kimi K3 Flagship Model Launches on Together AI(2026-07-28, 12 posts)
- Episode 14: Kimi K3 Max Tops Multiple Arena Leaderboards, Open-Source Model Rivals Proprietary(2026-07-28, 11 posts)
- Episode 15: Fireworks Test: Kimi K3 Matches Opus 5 Quality at Fraction of Cost(2026-07-28, 5 posts)
- Episode 16: Kimi K3 Impresses in Early Benchmarks, Sparking Buzz(2026-07-28, 3 posts)
- Episode 17: Local Kimi K3 Beats Cloud Models in 3D Physics Generation Test(2026-07-28, 5 posts)
- Episode 18: Deep Dive into Kimi K3 Tech Report: Engineering Synergy Drives State-of-the-Art Performance(2026-07-28, 21 posts)
- Episode 19: Moonshot AI's Kimi K3 Launches in Japan(2026-07-28, 2 posts)
- Episode 20: Kimi K3 Passes Compound Benchmark Amid Cost Efficiency Concerns(2026-07-29, 2 posts)
Primary sources
- Kimi K3 and Fable 5 now look much closer than the old open-vs-closed gap — FinanceYF5 · 2026-07-21
- Kimi K3 reaches 89.4% pass@4 and tops the benchmark over GPT-5.6 Sol — FinanceYF5 · 2026-07-21
- [source] Kimi K3 costs $4.65 per run and delivers 2.8× more work per dollar than Fable 5 — FinanceYF5 · 2026-07-21
- Kimi K3 leads on Go, but Fable 5 wins Python, JavaScript, TypeScript and Rust — FinanceYF5 · 2026-07-21
- [source] Kimi K3 hits 89.4% peak on software tasks while Fable 5 is slightly steadier — FinanceYF5 · 2026-07-21
- Kimi K3 and Fable 5 show nearly identical failure patterns on a software benchmark — FinanceYF5 · 2026-07-21
- DeepSWE Eval: Kimi K3 Matches Claude Fable 5 at 35% of the Cost — togethercompute · 2026-07-22