Kimi K3 Closes Gap with Fable 5 in Software Tasks, Signaling Open-Source Parity

Latest benchmark data shows that the open-source model Kimi K3 performs highly similarly to Fable 5 on software engineering tasks, boasting a task-by-task correlation coefficient of 0.72—a record high for models from different vendors. This indicates that frontier open-source models are no longer six months behind closed-source ones, signaling a substantial shift in the industry landscape.

Key Details and Performance Comparison

Looking at core pass metrics, Kimi K3 trails Fable 5 by just 1 point on pass@1. However, with a larger sampling budget, Kimi K3 begins to show an advantage, hitting 82.0% on pass@2 and 89.4% on pass@4, surpassing GPT-5.6 Sol. In terms of programming languages, Fable 5 takes the lead across Python, JavaScript, TypeScript, and Rust, while Kimi K3 overtakes it in Go (79 points vs. 71 points). Furthermore, their failure modes are nearly identical: about 65% of failures fall into the "almost succeeded" category, and both maintain their baselines well with regression rates of 11% and 10%, respectively. Because there are no extreme cases where one model consistently passes a task while the other consistently fails, the analysis suggests that this benchmark might be approaching saturation.

Compute Economics and Cost Advantages

While demonstrating equivalent performance, Kimi K3 possesses significant compute economic advantages. Data reveals that a single run of Kimi K3 costs only $4.65, whereas Fable 5 costs a hefty $13.41. Calculated per $100 invested, Kimi K3 can solve 14.7 tasks, which is 2.8 times the processing capacity of Fable 5.

2026-07-21 ~ 2026-07-21 · 6 related posts

Full story(20 episodes)→