Coding Agent Benchmark: Kimi Code Beats Claude Code and Codex
TheZachMueller · x · 2026-08-05
In a recent TerminalBench evaluation across 26 tasks, Kimi Code took the top spot with a 21/26 success rate, followed by Hermes Agent and Pi Agent. Claude Code and OpenAI Codex lagged behind, completing only 19 and 17 tasks respectively.
The benchmark revealed that Codex struggled with large workflows, hitting the 900-second limit twice on a reimbursement audit that Pi Agent finished in just 446 seconds. The author also pointed out that developers often overlook the fact that running models outside their official bundled environments (like Claude outside of Claude Code) might yield better performance or be more cost-effective.
Related event: Agent Framework Choice Massively Impacts LLM Costs(3 posts)→
More from coding & agent
- Google Open-Sources Gemini API Skills, Boosting Agent Code Generation to 96% — patloeber · 2026-08-05
- Malicious GitHub Repo Disguised as Crypto Exploit Exposed via LLM-Assisted Review — RSync25 · 2026-08-05
- AI Agents Automate Competitor Analysis and Influencer Marketing Strategy — fekdaoui · 2026-08-05
- Opinion: AI Coding is Manageable, but AI Workflows Risk Becoming Slop Without QA — oran_ge · 2026-08-05
- /human-review: Give AI Feedback Like Editing a Google Doc — petergyang · 2026-08-05
- Developer Recreates Pacman Entirely from AI-Generated Binary Code — Dimillian · 2026-08-05