Moore Threads says it pre-trained a 236B MoE model on a 10,000-GPU cluster

pstAsiatech · x · 2026-07-22

A quoted post claims that at last week’s World AI Conference in China, Moore Threads said a 236B-parameter MoE model was pre-trained from scratch on a domestic cluster with around 10,000 GPUs.

If true, the claim would put the training setup in the same league as the large Huawei-ecosystem clusters previously associated with that scale of model training.

Original post →

More from Infra

Infra channel →