LongCat-2 report says 1.6T model was trained entirely on domestic chips
ThePrimeClock · reddit · 2026-07-23
A Reddit thread dissects the LongCat-2 technical report and concludes that Meituan is publicly demonstrating a frontier-scale model trained and served entirely on domestic Chinese silicon, without Nvidia GPUs.
The post cites the report’s own Chinese meta description, which says the 1.6-trillion-parameter model was trained completely on domestic chips. It also highlights several infrastructure details: 50,000+ ASICs for pretraining, tens of thousands grouped into 48-chip “superpods,” RoCE networking between pods, and a two-tier design that reportedly lifts throughput by about 30%.
The thread also notes the software compensations needed for less memory per chip than an H800: ZeRO-1 sharding, recomputation, activation offloading, custom deterministic operators, binary-tree reduction to limit floating-point error, bit-flip detection, and fault isolation for bad network links.
More from Companies & People
- Zhipu chief says open source is the company’s core agenda, and restraint is key to survival — teortaxesTex · 2026-07-23
- Google’s Anthropic bet is a strange way to fund its future rival — bindureddy · 2026-07-23
- Kaggle says competitors are shifting from coding to directing agent teams — zakelfassi · 2026-07-23
- Artificial Analysis is hiring eval engineers after Amazon AGI layoffs — ArtificialAnlys · 2026-07-23
- Wei Shaojun’s background adds context to Oriental Computing’s DF1000 chip — teortaxesTex · 2026-07-23
- Amazon AGI team cuts jobs as researcher seeks computer-use agent roles — hamostaf04 · 2026-07-23