LongCat-2 report says 1.6T model was trained entirely on domestic chips

ThePrimeClock · reddit · 2026-07-23

A Reddit thread dissects the LongCat-2 technical report and concludes that Meituan is publicly demonstrating a frontier-scale model trained and served entirely on domestic Chinese silicon, without Nvidia GPUs.

The post cites the report’s own Chinese meta description, which says the 1.6-trillion-parameter model was trained completely on domestic chips. It also highlights several infrastructure details: 50,000+ ASICs for pretraining, tens of thousands grouped into 48-chip “superpods,” RoCE networking between pods, and a two-tier design that reportedly lifts throughput by about 30%.

The thread also notes the software compensations needed for less memory per chip than an H800: ZeRO-1 sharding, recomputation, activation offloading, custom deterministic operators, binary-tree reduction to limit floating-point error, bit-flip detection, and fault isolation for bad network links.

Original post →

More from Companies & People

Companies & People channel →