Ant's Ling-3.0-flash Activates Only 5.1B Params: Architecture and Cost Analysis
jkris050 · reddit · 2026-08-05
Ant Group's lab inclusionAI open-sourced the Ling-3.0-flash model under the MIT license on Aug 4. It has 124B total parameters but only activates 5.1B (about 1/64) per token, utilizing 512 routed experts and 1 shared expert.
Architectural Highlights:
- Natively adopted hybrid linear attention from the start of pretraining rather than retrofitting.
- Features 35 KDA layers alternating 5:1 against 7 gated MLA layers.
Performance & Cost Debate:
- The lab claims it matches or beats their own 1T param flagship, Ring-2.6-1T.
- The poster notes that if a lab can cut activated params by 12x without losing performance, it redefines the "cheap tier"—it's not a degraded version, but the same capability firing fewer experts.
- Deployment Pain Points: Currently lacks GGUF and llama.cpp support, requiring 4 GPUs for custom vLLM/SGLang forks. It's cheap to run, but holding 124B weights means VRAM ownership costs remain high.
More from Infra
- Astera Labs Predicts NPO Deployment in 2027, CPO to Follow in 2028 — bookwormengr · 2026-08-05
- US Chip Export Controls Backfire: Samsung and SK Hynix Turn to Chinese Toolmakers — kevinsxu · 2026-08-05
- TriAttention Integrated into TensorRT-LLM for Efficient Long-Context Inference — songhan_mit · 2026-08-05
- SpaceX Market Cap Drops $130B Overnight After $15.8B AI Spending Spree in Q2 — 智东西 · 2026-08-05
- Ex-OpenAI Exec Slams Goldman Sachs Token Demand Forecast, Cites 100x Cost Drop — ChrSzegedy · 2026-08-05
- SK Hynix and Samsung Evaluate AMEC Etchers for Chinese Fabs — zephyr_z9 · 2026-08-05