Six agent harnesses compared: token usage per task on identical model (Kimi K3)
zainhas · x · 2026-09-19
The author benchmarked input and output token usage across 6 different agent harnesses on 12 tasks, all solved/saturated, using the same model (Kimi K3) for every task to isolate harness differences. Findings: Claude Code is pretty token-hungry, Pi outputs a lot of tokens, and Codex sits in the middle of the pack. Lower is better across all tasks.
More from coding & agent
- EvoSkill v2 treats agent skills as executable state—and found agents learning to cheat — rohanpaul_ai · 2026-09-19
- Musecases Launches: A Community-Voted Prompt Library for AI Agents — ChrisUniverse · 2026-09-19
- ~40,000 passing tests: dev explains why he doesn't review every line of AI code — doodlestein · 2026-09-19
- NVIDIA's SoL-Pi GitHub repo: auto-research loops for efficient agent harnesses — aigclink · 2026-09-19
- Blogger: most engineers are just vibe coders, real skills lie in fundamentals — ashishllm · 2026-09-19
- Installers are going away — prompts with installer skills are next — BLUECOW009 · 2026-09-19