Coding Agent Benchmark: Kimi Code Beats Claude Code and Codex

TheZachMueller · x · 2026-08-05

In a recent TerminalBench evaluation across 26 tasks, Kimi Code took the top spot with a 21/26 success rate, followed by Hermes Agent and Pi Agent. Claude Code and OpenAI Codex lagged behind, completing only 19 and 17 tasks respectively.

The benchmark revealed that Codex struggled with large workflows, hitting the 900-second limit twice on a reimbursement audit that Pi Agent finished in just 446 seconds. The author also pointed out that developers often overlook the fact that running models outside their official bundled environments (like Claude outside of Claude Code) might yield better performance or be more cost-effective.

Related event: Agent Framework Choice Massively Impacts LLM Costs(3 posts)→

Original post →

More from coding & agent

coding & agent channel →