Tests Show Agent Harness Choice Can Triple LLM Costs

TheZachMueller · x · 2026-08-05

A developer points out that many users overlook model-harness compatibility when using AI coding agents. Running GPT or Claude outside their native environments like Codex or Claude Code might yield better or more cost-effective results.

Cited tests reveal that running Kimi K3 across six different harnesses on 26 identical tasks leads to a cost spread of up to 3x the average. Consequently, the team prioritized selecting a harness and open-sourced 300M distilled tool-calling tokens, urging the community to rethink the relationship between models and harnesses.

Related event: Agent Framework Choice Massively Impacts LLM Costs(3 posts)→

Original post →

More from coding & agent

coding & agent channel →