Composio benchmark finds similar success rates but wide speed gaps across coding harnesses
Teknium · x · 2026-07-29
A Composio comparison shows that three coding harnesses ended up with similar success rates, but very different speed and token efficiency.
- Kimi Code: 22/28 tasks
- Hermes: 21/28 tasks
- Claude Code: 20/28 tasks
The split was in runtime:
- Hermes: median 179s per task
- Kimi Code: 297s
- Claude Code: 348s
In short, the fastest harness and the most token-efficient harness were not the same one.
More from coding & agent
- Relume MCP and Claude Code power a Markdown-first design workflow — michalmalewicz · 2026-07-29
- A research talk compares SWE-bench, CodeClash, and ProgramBench for agentic coding — OfirPress · 2026-07-29
- Kuna debuts as an agent-driven Rust decompiler, ranking near Hex-Rays in structuring — moyix · 2026-07-29
- $100k Bet: AI 'Vibecoding' Can't Match Complex SEO Tools Like Ahrefs — eptwts · 2026-07-29
- An AI-agent startup argues billing should work from facts, not scattered API calls — davidcrawshaw · 2026-07-29
- Most Cursor users still treat it like VS Code, despite 1M users and a $29.3B valuation — nikola_mr64990 · 2026-07-29