Testing Kimi K3: Up to 30x Token Cost Difference Across Agent Harnesses

evijit · x · 2026-07-30

Composio evaluated Kimi K3 across three different agent harnesses (Claude Code, Hermes, and Kimi Code) using 28 identical tasks.

The results show that while all three harnesses achieved similar success rates in completing the tasks, there was a massive disparity in token efficiency. Depending on the harness used, the token cost for the exact same task varied by up to 30x.

Related event: Kimi K3 Tests Reveal Massive Token Consumption Gap Across Agent Frameworks(2 posts)→

Original post →

More from coding & agent

coding & agent channel →