Kimi K3 misses more softly, while GPT-5.6 Sol breaks baselines more often
zainhas · x · 2026-07-23
Kimi K3 and Sol fail in different ways
The attached analysis breaks down failure modes on the same benchmark:
- Kimi K3 gets closer to the target but more often stops short of passing every test.
- GPT-5.6 Sol breaks more baseline tests.
The chart quantifies the difference:
- Kimi K3: 65% near misses, 11% partial, 13% big misses, 11% broke baseline
- Sol: 54% near misses, 15% partial, 11% big misses, 20% broke baseline
The author notes Sol’s 20% baseline-breaking failures are consistent with other GPT models, while Kimi’s profile looks more like Claude-style regressions.
Related event: Kimi K3 Max vs GPT-5.6 Sol Max: A Routing Problem in Coding(10 posts)→
More from Models
- BullshitBench update: GPT-6-Astra beats all prior OpenAI models but still trails Anthropic — scaling01 · 2026-09-11
- Astra Scores 83% on GauntletBench, First Computer-Use Agent to Beat Human Baseline — ducha_aiki · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- DeepSeek V4 Pro API to continue after Sept 2026, billing unchanged — teortaxesTex · 2026-09-11
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11