Huawei's 100K-card Super Cluster can train a 10T-param model on 100T tokens in 30 days

teortaxesTex · x · 2026-09-18

Per @tphuang, Huawei's 100,000-card Super Cluster—built from 25 Atlas-960 SuperNodes—can train a 10-trillion-parameter model on 100T tokens in 30 days.

A significant datapoint in the race for hyperscale AI compute.

Original post →

More from Infra

Infra channel →