Ant Group launches Ling-3.0-flash with 124B parameters and a free API until Aug. 3
Loose_Bank1709 · reddit · 2026-07-27
Ant Group’s model team says it is releasing Ling-3.0-flash, a hybrid-reasoning MoE model aimed at production-scale agents.
The model has 124B total parameters, 5.1B active parameters per token, and a 256K context window. Ant claims it matches or beats its 1T flagship on most of the reported benchmarks, while using only 1/8 of the total and 1/12 of the active parameters.
The API is free until August 3, but the weights are not included: this is not an open-weights release, and there is no downloadable checkpoint. The same team notes that the previous Ling-2.6-flash was released under MIT, making this a very different distribution strategy.
More from Infra
- Sparrow switches its Standard mode to Ministral 3 14B for local document extraction — andrejusb · 2026-07-27
- Nvidia supplier Wistron opens $700 million Texas plant for GB300 and Vera Rubin systems — Beth_Kindig · 2026-07-27
- BeeLlama.cpp v0.4.1 adds KV-cache precision tails and new quantization modes — Anbeeld · 2026-07-27
- Running 100 million tokens through GLM 5.2 NVFP4 locally costs about $1 — _akhaliq · 2026-07-27
- Cloudflare’s AI-training block can also stop Googlebot after September 15 — daluoseo · 2026-07-27
- Self-hosted proxy unifies 424 AI models behind one OpenAI-compatible endpoint — ranadheer535 · 2026-07-27