Kimi K3 hits 89.4% peak on software tasks while Fable 5 is slightly steadier
FinanceYF5 · x · 2026-07-21
Kimi K3 and Fable 5 reach a per-task correlation of 0.72, the highest cross-vendor similarity reported here. The analysis says no task produced a perfect 4/4 split where one model always passed and the other always failed, suggesting the benchmark may be saturating.
Kimi K3 reaches a peak score of 89.4%, with 76.6% stable solve rate and 45 tasks solved all four times. Fable 5 is slightly more stable at 79.0%, with 58 tasks solved four times in a row. The accompanying plot shows Kimi has the widest overall reach in the dataset at 89.4% coverage.
More from Models
- Moonshot pauses Kimi K3 signups five days after launch as GPU demand surges — eyishazyer · 2026-07-21
- AI Diplomacy demo makes agents negotiate, ally, and betray each other — jamdac · 2026-07-21
- Newer models need a different prompting style, and old tricks can make outputs worse — emollick · 2026-07-21
- GLM-5.5 is said to arrive in 4 weeks with open weights — tanay_mehta · 2026-07-21
- Fable 5 is credited with a 3-variable counterexample to the Jacobian conjecture — Various-Affect4841 · 2026-07-21
- Ben’s Bites roundup highlights Kimi K3, Fable 5, Cursor costs and self-driving companies — Ben's Bites · 2026-07-21