Leak: GLM-5 Was a Modest 40B-Active, 28.5T-Token Pretrain, Insider Says
teortaxesTex · x · 2026-08-22
Community poster teortaxesTex reacted to speculation that GLM 5.3 was delayed 2 months with a possible agentic-specialized "5.3 Turbo" variant, arguing it's not just about agentic specialization.
He claims (unverified) that GLM-5's pretraining was notably unambitious: off-the-shelf DSA, 40B active parameters, 28.5T tokens — mid-pack but good enough, and suggests Zhipu may have improved pretraining skills since.
More from Models
- Researcher: GPT 5.6 Sol Ultra Beats Pro for Long-Horizon Hard Problems — arankomatsuzaki · 2026-08-24
- Google Criticized: Gemini 3.7 Still Missing From Its Own Jules Agent a Week Later — brandon_galang · 2026-08-24
- Qwen 27B 3.8 low quantization tested: Q3 XXS works well locally — jeremyckahn · 2026-08-24
- Users notice significant quality shift in GPT-5.6 output — haider1 · 2026-08-24
- Ramp Stats: Anthropic Opus 4.8 and Sonnet 4.6 Lead Usage — vista8 · 2026-08-24
- Tencent Releases UI-Mate-27B, a Desktop GUI Agent Model — tencent · 2026-08-24