Moore Threads says it pre-trained a 236B MoE model on a 10,000-GPU cluster
pstAsiatech · x · 2026-07-22
A quoted post claims that at last week’s World AI Conference in China, Moore Threads said a 236B-parameter MoE model was pre-trained from scratch on a domestic cluster with around 10,000 GPUs.
If true, the claim would put the training setup in the same league as the large Huawei-ecosystem clusters previously associated with that scale of model training.
More from Infra
- Will AI agents really need crypto wallets, or just better payment APIs? — Digitalpaver · 2026-07-22
- After GLM goes live, the next test is DSv4 Flash on 8×H100 — TheZachMueller · 2026-07-22
- llama.cpp adds support for Laguna XS.2 and M.1 in release b10087 — LaurentPayot · 2026-07-22
- LightOn Search hits 5th place on BrowseComp-Plus with 86.27% accuracy — CShorten30 · 2026-07-22
- Nvidia’s Spectrum switches add CPO support as Lambda tests early units — TheZachMueller · 2026-07-22
- Hermes Cloud adds on-demand disk expansion and CPU/RAM upgrades — Teknium · 2026-07-22