Kimi K3 Tops pmpp-hard Eval: Open Model Achieves SOTA in Hardcore Agent Tasks
xeophon · x · 2026-08-03
In the newly released pmpp-hard evaluation, open-source models demonstrated impressive agent capabilities. The benchmark involved 69 GPU kernel tasks, conducting 3,100 agent rollouts across 11 models and consuming over 5.8 billion tokens. Kimi K3 ultimately took first place with a score of 0.71, proving that open models can achieve SOTA performance even when handling complex and hardcore tasks.
More from coding & agent
- Current AI Agents Lack the Human 'Power Law of Practice' — xwang_lk · 2026-08-03
- 95% of agent skills missing activation directives, GitHub scan reveals — rseroter · 2026-08-03
- Tracking 'Shadow AI' in WordPress: Monitoring Unauthorized HTTP Calls — HaktanSuren · 2026-08-03
- Microsoft Launches IQ Live Series for Enterprise Context-Aware Agent Architectures — DanWahlin · 2026-08-03
- Tech Giants Build Context Layers: Ontologies and Knowledge Graphs Emerge as AI's New Battleground — brucemacv · 2026-08-03
- Out of Codex Credits? Developers Discuss Pivoting to OpenCode and Alternatives — yacineMTB · 2026-08-03