Kimi K3 Leads Multi-Sampling Benchmarks
zephyr_z9 · x · 2026-07-18
The image compares the performance of **Kimi K3** and **Fable 5** across different sampling attempts: - `pass@1`: Kimi K3 scores 68.5, slightly below Fable 5's 69.9. - `pass@2`: Kimi K3 scores 82.0, beating Fable 5's 80.2. - `pass@4`: Kimi K3 scores 89.4, ahead of Fable 5's 88.5. The post highlights that Kimi K3 performs significantly better under multi-sampling settings, earning it the title of "open/closed SOTA".
More from Models
- Claude 20x users report sharply tighter limits and faster quota burn — MarcJSchmidt · 2026-07-21
- Cola launches July, the latest model jokingly billed as “second only to Fable” — oran_ge · 2026-07-21
- Kimi K3 looks stronger and about 5× cheaper on a frontend dashboard task — OwariDa · 2026-07-21
- Last Week in AI recap: Anthropic’s $65B round, IPO filing, and Microsoft’s MAI push — Last Week in AI · 2026-07-21
- A user says Claude 4.6 felt worse yesterday and asks whether model quality can drift over time — Rahios · 2026-07-21
- Kimi K3 hits 89.4% peak on software tasks while Fable 5 is slightly steadier — FinanceYF5 · 2026-07-21