How Meituan's LongCat lab trained a frontier model on Huawei 910C chips
bookwormengr · x · 2026-09-21
A deep dive into the rise of Meituan's LongCat AI Lab: LongCat 2.0 is the first publicly known model trained on Huawei 910Cs, most likely on CloudMatrix 384 Superpods (384 NPUs each, vs 72 in Nvidia NVL72 racks), with the blog claiming 50K total ASICs used.
The author argues the community noticed the model but missed the engineering ingenuity: overcoming US export controls with 'weaker' components, reduced HBM dependence, and new architectures. LongCat mastered hard-to-train sparse MoE, invented improved attention, and independently developed N-gram-based sparsity as a new sparsity axis. The piece frames it as China's AI ecosystem working together in ways unthinkable for Western peers like Uber or DoorDash.
More from Infra
- vLLM ships Hybrid KV Cache Manager for mixed-attention model inference — TheZachMueller · 2026-09-22
- SGLang's hicache: use an L3 storage cache to keep KV cache alive across local model swaps — TheZachMueller · 2026-09-22
- Egypt's AI Ecosystem Hits Production Scale With $400M Data Center, 10x NVIDIA Learner Growth — nordicinst · 2026-09-22
- RTX Pro 6000 vs a used 3090 vs cloud rental: the LoRA training math — big-in-jap · 2026-09-21
- ComfyUI GPU rental showdown: Modal's 35s cold starts and free 1TiB beat RunPod — ronalder100 · 2026-09-21
- Meta partners with Arm on Arm AGI CPU, its first AI-era data center CPU — bookwormengr · 2026-09-21