Kimi K3 Tops pmpp-hard Eval: Open Model Achieves SOTA in Hardcore Agent Tasks

xeophon · x · 2026-08-03

In the newly released pmpp-hard evaluation, open-source models demonstrated impressive agent capabilities. The benchmark involved 69 GPU kernel tasks, conducting 3,100 agent rollouts across 11 models and consuming over 5.8 billion tokens. Kimi K3 ultimately took first place with a score of 0.71, proving that open models can achieve SOTA performance even when handling complex and hardcore tasks.

Original post →

More from coding & agent

coding & agent channel →