Kimi K3 Max trails on pass@1 but beats Sol Max at pass@4 on SWE tasks
zainhas · x · 2026-07-23
A deep dive compares Kimi K3 Max with Sol Max on software-engineering tasks and argues that higher variability can translate into a better pass@k ceiling.
Key numbers from the chart:
- pass@1: Sol 72.7% vs Kimi 68.5%
- pass@2: Kimi 82.0% vs Sol 81.0%
- pass@4: Kimi 89.4% vs Sol 85.8%
The author’s takeaway is that Kimi’s higher variance hurts single-shot performance, but gives it a stronger portfolio outcome once multiple attempts are allowed.
Related event: Kimi K3 Max vs GPT-5.6 Sol Max: A Routing Problem in Coding(10 posts)→
More from Models
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- DeepSeek V4 Pro API to continue after Sept 2026, billing unchanged — teortaxesTex · 2026-09-11
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11
- TheZvi Polls: Has Your Coding Model Choice Changed Since Fable 5.1 and Astra? — TheZvi · 2026-09-11
- antirez Weighs In on Anthropic Banning Minors From Using Claude — antirez · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11