Kimi K3 Max trails on pass@1 but beats Sol Max at pass@4 on SWE tasks

zainhas · x · 2026-07-23

A deep dive compares Kimi K3 Max with Sol Max on software-engineering tasks and argues that higher variability can translate into a better pass@k ceiling.

Key numbers from the chart:

The author’s takeaway is that Kimi’s higher variance hurts single-shot performance, but gives it a stronger portfolio outcome once multiple attempts are allowed.

Related event: Kimi K3 Max vs GPT-5.6 Sol Max: Complementarity and Routing in Coding(8 posts)→

Original post →

More from Models

Models channel →