Kimi K3 Max trails on pass@1 but beats Sol Max at pass@4 on SWE tasks
zainhas · x · 2026-07-23
A deep dive compares Kimi K3 Max with Sol Max on software-engineering tasks and argues that higher variability can translate into a better pass@k ceiling.
Key numbers from the chart:
- pass@1: Sol 72.7% vs Kimi 68.5%
- pass@2: Kimi 82.0% vs Sol 81.0%
- pass@4: Kimi 89.4% vs Sol 85.8%
The author’s takeaway is that Kimi’s higher variance hurts single-shot performance, but gives it a stronger portfolio outcome once multiple attempts are allowed.
Related event: Kimi K3 Max vs GPT-5.6 Sol Max: Complementarity and Routing in Coding(8 posts)→
More from Models
- GPT 5.6 Sol and Claude refused to game an AI detector, while Grok 4.5 passed after 14 drafts — aman_madaan · 2026-07-23
- Muse Spark 1.1 generated 3D arcade games for $0.026, versus $0.073 for Gemini 3.6 — rohanpaul_ai · 2026-07-23
- Tesla Robotaxi fleet is already testing early FSD v15 builds, post claims — ChrisGPT · 2026-07-23
- OpenAI GPT-5.6 Sol helped solve 6 open Erdős problems in 5 days — Charuru · 2026-07-23
- AI designer says blocking Reddit training data could make future LLMs better — AIandDesign · 2026-07-23
- Google says Gemini is now routed into Search AI Overviews and AI Mode — gaganghotra_ · 2026-07-23