Kimi K3 looks strong on benchmarks, but a real coding test says reliability still lags
Cole Medin · youtube · 2026-07-24
A YouTube creator benchmarks Kimi K3 against Opus 4.8 and Kimi K2.7 on real engineering tasks and argues the public hype overstates what the model can actually do.
Key takeaways from the video:
- Kimi K3 looks extremely strong on benchmarks and may appear close to or ahead of top closed models.
- In practical engineering work, it still shows reliability failure modes that public leaderboards don’t capture.
- On simple tasks, K3 can roughly tie Opus; on more complex tasks, Opus pulls ahead.
- The creator says open-weight models are still valuable as the workhorse in a mixed-model workflow.
- The video includes an open-source benchmark repo, rubric, prompts, and harness so others can reproduce the tests.
A notable result cited in the video is a failure-rate gap of 8% vs. 36%, used to argue that benchmark design matters more than leaderboard headlines.
More from Models
- Repligate says Claude Opus 3 appears to evolve without changing its weights — repligate · 2026-07-27
- “Opus 5” post lands as a rebenchmarking-at-scale AI joke — kalomaze · 2026-07-27
- Top models now write worse than a year ago, critic says — dbreunig · 2026-07-27
- MPT-30B radar charts became an unexpectedly controversial design choice — code_star · 2026-07-27
- Local Gemma 4 31B starts acting sarcastic and users cannot reproduce it — n0head_r · 2026-07-27
- Google’s Gemini 3.6 Flash could win by matching Sonnet quality at a lower cost — haider1 · 2026-07-27