New Season of LLMs: Kimi K3 Leads, GLM5.3 Extends, Chinese Models in Fierce Competition
葬AI · wechat · 2026-08-17
A WeChat article provides deep analysis of the Chinese LLM competitive landscape in H1 2026. The author argues public benchmarks are meaningless as vendors can produce any scores; real-world performance matters. The new season has two phases: first half focuses on coding, with Zhipu GLM5.2 defeating Opus4.8; second half sees parameter scaling, with Kimi K3 (2T) leading, approaching Fable5 scores and excelling in front-end generation. Qwen3.8Max (2.4T) follows but lacks multimodal strength. DeepSeek V4 Pro is unstable, with official benchmarks deviating from real tests. Grok4.6 offers good value, and Musk hints at Grok4.7 (2.1T). GLM5.3 update shows limited improvement and slower execution. The author concludes Kimi is the version leader, while others must prove themselves at 2T scale.
More from Companies & People
- Developer habits shift: Agents become collaborators from simple tools — latticecut · 2026-08-24
- SoftBank planning record $6.3B bond sale to fund OpenAI investments — tszzl · 2026-08-24
- Google Criticized: Gemini 3.7 Still Missing From Its Own Jules Agent a Week Later — brandon_galang · 2026-08-24
- Product Success Breeds Clones: What Is the Real Moat? — talkaboutdesign · 2026-08-24
- Enterprise AI budgets drained with no ROI: Where's the problem? — annetgriffin · 2026-08-24
- Meme of Anthropic employees rolling to the office post-IPO — zainhas · 2026-08-24