Chinese AI Models Scale Up: GLM 5.3 Flash Trains on 30T Tokens
teortaxesTex · x · 2026-08-28
Leaked details reveal progress in Chinese large-scale pre-training: GLM 5.3 Flash is an 18B active MoE architecture pretrained on 30T tokens, while OpenPangu 2.0 Pro boasts 505B total parameters (18B active) trained on Ascend chips. Other models like MooreThreads (236B MoE) and LongCat 2.0 (1600B total) suggest compute barriers are lowering.
More from Infra
- Designing a Reliable LLM Gateway: Building a Stable Backend Entry Point — Mahmoud_Zalt · 2026-08-28
- LightGlue ONNX: Local Feature Matching with TensorRT Support — rsasaki0109 · 2026-08-28
- LLM escapes VM three times, debate renews on sandbox definitions — dyn___ · 2026-08-28
- Cloudflare Engineer on Building Extensible Software for the LLM Era — 新智元 · 2026-08-28
- Google, NVIDIA Back Open Source Project to Optimize LLM Inference on Kubernetes — SumitGup · 2026-08-28
- Australia Minister: No Fossil Fuel Carve-out for Datacenters — nordicinst · 2026-08-28