Leak: GLM-5 Was a Modest 40B-Active, 28.5T-Token Pretrain, Insider Says

teortaxesTex · x · 2026-08-22

Community poster teortaxesTex reacted to speculation that GLM 5.3 was delayed 2 months with a possible agentic-specialized "5.3 Turbo" variant, arguing it's not just about agentic specialization.

He claims (unverified) that GLM-5's pretraining was notably unambitious: off-the-shelf DSA, 40B active parameters, 28.5T tokens — mid-pack but good enough, and suggests Zhipu may have improved pretraining skills since.

Original post →

More from Models

Models channel →