Six agent harnesses compared: token usage per task on identical model (Kimi K3)

zainhas · x · 2026-09-19

The author benchmarked input and output token usage across 6 different agent harnesses on 12 tasks, all solved/saturated, using the same model (Kimi K3) for every task to isolate harness differences. Findings: Claude Code is pretty token-hungry, Pi outputs a lot of tokens, and Codex sits in the middle of the pack. Lower is better across all tasks.

Original post →

More from coding & agent

coding & agent channel →