Kimi K3 hits 89.4% peak on software tasks while Fable 5 is slightly steadier
FinanceYF5 · x · 2026-07-21
Kimi K3 and Fable 5 reach a per-task correlation of 0.72, the highest cross-vendor similarity reported here. The analysis says no task produced a perfect 4/4 split where one model always passed and the other always failed, suggesting the benchmark may be saturating.
Kimi K3 reaches a peak score of 89.4%, with 76.6% stable solve rate and 45 tasks solved all four times. Fable 5 is slightly more stable at 79.0%, with 58 tasks solved four times in a row. The accompanying plot shows Kimi has the widest overall reach in the dataset at 89.4% coverage.
Related event: Kimi K3 Matches Fable 5 in SWE Benchmarks at a Third of the Cost(7 posts)→
More from Models
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- DeepSeek V4 Pro API to continue after Sept 2026, billing unchanged — teortaxesTex · 2026-09-11
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11
- TheZvi Polls: Has Your Coding Model Choice Changed Since Fable 5.1 and Astra? — TheZvi · 2026-09-11
- antirez Weighs In on Anthropic Banning Minors From Using Claude — antirez · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11