Kimi K3 finds 32 of 86 vulnerabilities for just $19.37 in benchmark run

zeeg · x · 2026-07-23

Kimi K3 tops Warden-style benchmark on cost efficiency

A post claims Kimi K3 performed reliably on Warden’s benchmark, with no failures and no repair jobs during the run.

The attached chart compares models on a vulnerability-finding task:

The poster says Kimi’s accuracy-to-price ratio is “amazing” and that they’re switching to it.

Original post →

More from Models

Models channel →