LongCat-2 report says 1.6T model was trained entirely on domestic chips
ThePrimeClock · reddit · 2026-07-23
A Reddit thread dissects the LongCat-2 technical report and concludes that Meituan is publicly demonstrating a frontier-scale model trained and served entirely on domestic Chinese silicon, without Nvidia GPUs.
The post cites the report’s own Chinese meta description, which says the 1.6-trillion-parameter model was trained completely on domestic chips. It also highlights several infrastructure details: 50,000+ ASICs for pretraining, tens of thousands grouped into 48-chip “superpods,” RoCE networking between pods, and a two-tier design that reportedly lifts throughput by about 30%.
The thread also notes the software compensations needed for less memory per chip than an H800: ZeRO-1 sharding, recomputation, activation offloading, custom deterministic operators, binary-tree reduction to limit floating-point error, bit-flip detection, and fault isolation for bad network links.
More from Companies & People
- Anthropic Insiders: Not Everyone at the Lab Believes in High p(doom) — anpaure · 2026-09-11
- Lumara AI Film Festival Comes to NYC Oct 26, Top AI Filmmakers to Compete — 0xAllen_ · 2026-09-11
- X drama: Anthropic researchers accused of spying on academic customers and racing them to results — basedjensen · 2026-09-11
- Investor argues Palantir-Nvidia partnership should slash Anthropic's IPO valuation — pdamodaran · 2026-09-11
- a16z podcast: why 2-3 person startups are absent from policy debates — a16z Podcast · 2026-09-11
- PyTorch Day Korea 2026 launches first offline conf, CFP closes Sept 13 — PyTorch · 2026-09-11