Harness matters: token-level study shows GLM-5.3, Inkling and Kimi-K3 behave differently across agent frameworks
bhutanisanyam1 · x · 2026-09-23
As part of the Reflection AI × Scale AI release, the author measured how harnesses change frontier model behavior:
- Ran GLM-5.3, Inkling and Kimi-K3 through mini, Pi and OpenCode, analyzing token-level differences;
- Kimi is the largest model but most token-efficient paired with Pi; GLM prefers running tests, Inkling spends most time reading files;
- Counterintuitively, task failures aren't from laziness — models spend twice the compute and steps trying harder, and failures stem from tool/harness mismatches.
Takeaway: pick your harness-model combination carefully.
More from Models
- Opus 5.5 launch met with shrug as enthusiasts shift to open models like Qwen and DeepSeek — ByteSize_Chaos · 2026-09-23
- Altman: GPT-6 Sol and Luna have no competition on per-task pricing — sama · 2026-09-23
- Gradio distills Qwen's 9B prompt rewriter into an 812MB 0.8B model that fits a laptop — Gradio · 2026-09-23
- Polymarket: GPT-6 claims up to 93% cheaper coding task costs than Claude Opus 5 — Polymarket · 2026-09-23
- OpenAI launches GPT-6 Sol and Luna, permanently cuts API prices 50% — aziz4ai · 2026-09-23
- Sam Altman announces GPT-6 Sol and Luna: major upgrades at half the price — sama · 2026-09-23