Kimi K3 and Fable 5 show nearly identical failure patterns on a software benchmark
FinanceYF5 · x · 2026-07-21
A comparison chart shows Kimi K3 and Fable 5 failing in very similar ways on a software engineering benchmark. About 65% of failures are near-misses for both models, and both largely preserve existing baselines, with regression rates of 11% for Kimi and 10% for Fable.
The accompanying analysis says the two models have a per-task correlation of 0.72, the highest cross-vendor similarity on record. It also notes that no task showed a full 4/4 success vs 0/4 failure split in either direction, which may indicate the benchmark is nearing saturation.
Related event: Kimi K3 Matches Fable 5 in SWE Benchmarks at a Third of the Cost(7 posts)→
More from Models
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11
- TheZvi Polls: Has Your Coding Model Choice Changed Since Fable 5.1 and Astra? — TheZvi · 2026-09-11
- antirez Weighs In on Anthropic Banning Minors From Using Claude — antirez · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- Fully local voice assistant on an RTX 3060 replicates the GPT Live demo in 6.5 minutes — liampetti · 2026-09-11
- Meta's Muse Agent has built-in invite code logic, hinting at free-usage expansion — testingcatalog · 2026-09-11