Open-weight models vs Claude Sonnet 5: GLM 5.3 wins 20 real coding tasks at 1/10th the cost
shensi · x · 2026-09-15
The mergeapi team benchmarked five open-weight models — GLM 5.3, GLM 5.3 Flash, DeepSeek V4 Pro, DeepSeek V4 Flash, and Kimi K3 — against Claude Sonnet 5 on 20 real coding tasks. GLM 5.3 came out on top while costing about one-tenth of Claude's price. Full results in the linked writeup.
More from coding & agent
- tldraw's AI design sprint: Codex prototypes every discussed interaction overnight, artifacts by day 2 — max__drake · 2026-09-15
- Research agents log what they cite — should they also defend what they reject? — Entire_Mark8010 · 2026-09-15
- Grok Bot Auto-Schedules Pinterest Posts, Writes Titles From Images—Where Claude Failed — prasenx · 2026-09-15
- Amplitude tripled PR volume in six months by fixing CI, not agents — mobileraj · 2026-09-15
- Looking for a minimal, near-instant CLI coding agent for one-off bash tasks — funbike · 2026-09-15
- LangChain's Chase on why agent memory never sticks: deciding what to remember is app-specific — blaizedsouza · 2026-09-15