Kimi K3 and Sol 4.6 show low correlation and cover 95.6% together
zainhas · x · 2026-07-23
The author says Kimi K3 and Sol 4.6 fail in different ways, with a low per-task correlation of 0.46.
- The benchmark suggests the two models complement each other rather than overlap heavily.
- A routing or cascading setup covers 108 of 113 tasks.
- That combination is described as the best two-model portfolio on the benchmark, reaching 95.6% task coverage.
Related event: Kimi K3 Max vs GPT-5.6 Sol Max: A Routing Problem in Coding(10 posts)→
More from Models
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11