Kimi K3 and Fable 5 show nearly identical failure patterns on a software benchmark

FinanceYF5 · x · 2026-07-21

A comparison chart shows Kimi K3 and Fable 5 failing in very similar ways on a software engineering benchmark. About 65% of failures are near-misses for both models, and both largely preserve existing baselines, with regression rates of 11% for Kimi and 10% for Fable.

The accompanying analysis says the two models have a per-task correlation of 0.72, the highest cross-vendor similarity on record. It also notes that no task showed a full 4/4 success vs 0/4 failure split in either direction, which may indicate the benchmark is nearing saturation.

Related event: Kimi K3 Closes Gap with Fable 5 in Software Tasks, Signaling Open-Source Parity(6 posts)→

Original post →

More from Models

Models channel →