Chinese AI Models Scale Up: GLM 5.3 Flash Trains on 30T Tokens

teortaxesTex · x · 2026-08-28

Leaked details reveal progress in Chinese large-scale pre-training: GLM 5.3 Flash is an 18B active MoE architecture pretrained on 30T tokens, while OpenPangu 2.0 Pro boasts 505B total parameters (18B active) trained on Ascend chips. Other models like MooreThreads (236B MoE) and LongCat 2.0 (1600B total) suggest compute barriers are lowering.

Original post →

More from Infra

Infra channel →