Dev Warns: Don't Trust Benchmarks, Kimi Wins in Real-World Code Reviews

JensHonack · x · 2026-07-18

A developer points out that in practical scenarios like code reviews, blindly trusting official benchmarks can be misleading: a model rated "high" actually underperformed compared to another rated "low".

The post uses this to poke fun at the situation, expressing hope that the upcoming Kimi 3 will also survive the ultimate "don't trust benchmarks" real-world test.

Original post →

More from coding & agent

coding & agent channel →