Coding Agent Benchmark: Kimi Code Beats Claude Code and Codex
TheZachMueller · x · 2026-08-05
In a recent TerminalBench evaluation across 26 tasks, Kimi Code took the top spot with a 21/26 success rate, followed by Hermes Agent and Pi Agent. Claude Code and OpenAI Codex lagged behind, completing only 19 and 17 tasks respectively.
The benchmark revealed that Codex struggled with large workflows, hitting the 900-second limit twice on a reimbursement audit that Pi Agent finished in just 446 seconds. The author also pointed out that developers often overlook the fact that running models outside their official bundled environments (like Claude outside of Claude Code) might yield better performance or be more cost-effective.
Related event: Agent Framework Choice Massively Impacts LLM Costs(3 posts)→
More from coding & agent
- Agent Substrate Runtime Can Suspend and Resume Per Tool Call — jonathangrahl · 2026-09-21
- Turn any local LLM into a confidence-scored classifier via logprobs, full llama.cpp recipe included — DivideHorror3217 · 2026-09-21
- Ruff author charliermarsh: he only started using agents meaningfully in December 2025 — charliermarsh · 2026-09-21
- Open-source LLaMA-Factory fine-tunes 100+ LLMs; 200 examples can beat frontier models — Roger_M_Taylor · 2026-09-21
- SWE-2 free across Devin Cloud Agents, CLI and Desktop until October 8 — silasalberti · 2026-09-21
- Delta and OpenTable block AI agents, hinting at a coming platform stand-off — Scobleizer · 2026-09-21