Chinese coding models use 2-3x the tokens, so cheaper per token isn't cheaper overall
craigbalding · x · 2026-09-18
Based on coding benchmarks, the author observes Chinese open-source models tend to consume 2-3x the tokens on coding tasks — faster/cheaper per token doesn't mean lower total cost. They are materially stronger in terminal benchmarks with fast time-to-first-token, but are 'token drunkards'. Off-peak pricing from some providers (e.g. BitDeer) can halve costs and make them contenders. Since benchmarks can mislead, test your own workloads; no compelling reason to switch yet.
More from coding & agent
- Abandoned AI agent projects are uniquely hard to hand off, evaluator observes — Mariav_Dowdf · 2026-09-18
- PCP v0.1: A Draft Protocol for Purpose-Bound, Revocable AI Agent Authority — sierracatalina · 2026-09-18
- Legatus v0.1: open-source transport-independent coordination protocol for delegated agent work — sierracatalina · 2026-09-18
- Salesforce launches Trusted Enterprise AI Harness to unify agent context, governance and security — emmanuelvivier · 2026-09-18
- Running Codex, Claude and Pi Agents Safely: gVisor Sandboxes Plus tart macOS VMs — craigbalding · 2026-09-18
- EvalSeal: open-source tool shows LLM judges flip verdicts on 5 of 20 borderline eval cases — Fit_Fortune953 · 2026-09-18