Same model, 9.2x token spread: harness matters more than the model
zainhas · x · 2026-09-08
Real-world testing shows the agent harness can impact token usage and end-to-end time more than the model itself. Running GLM 5.3 Flash Max across three frameworks: Codex used 475K tokens in 9 mins, OMP 1.74M in 30 mins, and Opencode 4.36M in 20 mins — a 9.2x token spread on one model. Takeaway: choosing the right harness matters as much as choosing the model for cost and efficiency.
Related event: Coding Harness Matters More Than Model: Token Use Varies 9.2x(2 posts)→
More from coding & agent
- Autonomous NoSpoon microdrama agent gets major upgrade, public release soon — Kyrannio · 2026-09-08
- AI slop creates an order of magnitude more tech debt and kills the will to clean it up — zetalyrae · 2026-09-08
- AstraBlender lets ChatGPT drive cloud Blender from your phone, open-sourced — Short-Patient7772 · 2026-09-08
- AI forum grows to 320 personas, loses omniscience, adds 1-on-1 chat — mrjeeves · 2026-09-08
- Physicist's $100 Bounty Unclaimed for 20 Years Finally Solved by an AI Agent — IgorCarron · 2026-09-08
- Dev finds Astra falls short of Fable 5.1 on coding, needs extra review turns — bindureddy · 2026-09-08