Tests Show Agent Harness Choice Can Triple LLM Costs
TheZachMueller · x · 2026-08-05
A developer points out that many users overlook model-harness compatibility when using AI coding agents. Running GPT or Claude outside their native environments like Codex or Claude Code might yield better or more cost-effective results.
Cited tests reveal that running Kimi K3 across six different harnesses on 26 identical tasks leads to a cost spread of up to 3x the average. Consequently, the team prioritized selecting a harness and open-sourced 300M distilled tool-calling tokens, urging the community to rethink the relationship between models and harnesses.
Related event: Agent Framework Choice Massively Impacts LLM Costs(3 posts)→
More from coding & agent
- Agent Substrate Runtime Can Suspend and Resume Per Tool Call — jonathangrahl · 2026-09-21
- Turn any local LLM into a confidence-scored classifier via logprobs, full llama.cpp recipe included — DivideHorror3217 · 2026-09-21
- Ruff author charliermarsh: he only started using agents meaningfully in December 2025 — charliermarsh · 2026-09-21
- Open-source LLaMA-Factory fine-tunes 100+ LLMs; 200 examples can beat frontier models — Roger_M_Taylor · 2026-09-21
- SWE-2 free across Devin Cloud Agents, CLI and Desktop until October 8 — silasalberti · 2026-09-21
- Delta and OpenTable block AI agents, hinting at a coming platform stand-off — Scobleizer · 2026-09-21