Kimi K3 looks strong on benchmarks, but a real coding test says reliability still lags

Cole Medin · youtube · 2026-07-24

A YouTube creator benchmarks Kimi K3 against Opus 4.8 and Kimi K2.7 on real engineering tasks and argues the public hype overstates what the model can actually do.

Key takeaways from the video:

A notable result cited in the video is a failure-rate gap of 8% vs. 36%, used to argue that benchmark design matters more than leaderboard headlines.

Original post →

More from Models

Models channel →