Chinese coding models use 2-3x the tokens, so cheaper per token isn't cheaper overall

craigbalding · x · 2026-09-18

Based on coding benchmarks, the author observes Chinese open-source models tend to consume 2-3x the tokens on coding tasks — faster/cheaper per token doesn't mean lower total cost. They are materially stronger in terminal benchmarks with fast time-to-first-token, but are 'token drunkards'. Off-peak pricing from some providers (e.g. BitDeer) can halve costs and make them contenders. Since benchmarks can mislead, test your own workloads; no compelling reason to switch yet.

Original post →

More from coding & agent

coding & agent channel →