Same model, 9.2x token spread: harness matters more than the model

zainhas · x · 2026-09-08

Real-world testing shows the agent harness can impact token usage and end-to-end time more than the model itself. Running GLM 5.3 Flash Max across three frameworks: Codex used 475K tokens in 9 mins, OMP 1.74M in 30 mins, and Opencode 4.36M in 20 mins — a 9.2x token spread on one model. Takeaway: choosing the right harness matters as much as choosing the model for cost and efficiency.

Related event: Coding Harness Matters More Than Model: Token Use Varies 9.2x(2 posts)→

Original post →

More from coding & agent

coding & agent channel →