Ant Group Releases Ling-3.0-flash: 124B Parameters with Just 5.1B Active
SonglinYang4 · x · 2026-07-25
Ant Group's Ling team has released Ling-3.0-flash, a hybrid-reasoning MoE model built for production-scale agents. The model has 124B total parameters but activates only 5.1B per token. With 1/8 of the total and 1/12 of the active parameters, it reportedly matches or beats their 1T flagship model on most benchmarks. The vLLM team also praised its announce-first, open-source-next approach, noting it provides a stable window for open-source inference projects to prepare for day-0 support.
Related event: Ant Group Releases Ling-3.0-flash MoE Model(12 posts)→
More from Models
- LiquidAI's LFM2.5-2.6B Hits 82 tok/s Decode on Mac with 128K Context — helloiamleonie · 2026-08-05
- Researcher Notes 'Procrastination' in 5p6 Sol Extra High Model — RylanSchaeffer · 2026-08-05
- Stress-Testing DeepSeek Subscriptions: Does It Beat Gemini Flash in Value? — teortaxesTex · 2026-08-05
- Best Local LLMs for Coding on a 128GB Mac? — Electronic_Back1502 · 2026-08-05
- Open Source AI Hits a Wall in Long-Running Agentic Loops — bindureddy · 2026-08-05
- Dev: I'd rather iterate 10 times with Gemini 3.6 Flash than wait 2 hours with Qwen 3.8 max — DynamicWebPaige · 2026-08-05