Cline says open-weight models work better when the harness lets them verify more
cpaik · x · 2026-08-04
Cline says its internal evaluations suggest many open-weight models such as DeepSeek, GLM, and Kimi are RL-trained to spend more tokens verifying their work.
The key claim is that these models often:
- run tests
- check builds
- re-read diffs before finishing
Cline argues that its harness performs better because it lets the model work the way it was trained to work, instead of pushing for minimal-token efficiency. At open-weight pricing, that extra verification can translate into better results for less money, and its benchmark runs reportedly show about a 20% gain.
More from coding & agent
- GitHub Stacked PRs repo shows how to split one big review into layered changes — DanWahlin · 2026-08-04
- Evedev is being recommended as the default framework for internal agents — cramforce · 2026-08-04
- Agent swarms fail when they optimize for process instead of shipping code — doodlestein · 2026-08-04
- A month-by-month meme tracks how coding-agent habits keep changing — unixterminal · 2026-08-04
- A critique says OpenCode’s agent design breaks KV cache and weakens security — JFPuget · 2026-08-04
- YC-backed Buildbox launches agent analytics for real user outcomes, not just evals — ycombinator · 2026-08-04