Moore Threads Pretrains 236B MoE on 10,000-GPU Cluster
Moore Threads announced at the WAIC that it pre-trained a 236-billion-parameter MoE model from scratch on a 10,000-GPU domestic cluster, highlighting rapid advancements in China's AI hardware and compute capabilities.
2026-07-22 ~ 2026-07-22 · 2 related posts
- Moore Threads says it pre-trained a 236B MoE model on a 10,000-GPU cluster — pstAsiatech · 2026-07-22
- China’s AI hardware scene now has a 10K-GPU cluster and a 236B MoE model — teortaxesTex · 2026-07-22