Ant Group’s InclusionAI launches Ling-3.0-flash as a cheap fast execution model
AcanthisittaOk1699 · reddit · 2026-07-25
A Reddit post highlights Ant Group’s InclusionAI shipping a new cheap-fast execution model, Ling-3.0-flash, and says it is free for a week.
- The model is described as a sparse MoE with 124B total parameters / 5.1B active, 256K context, and sub-100ms first token latency.
- The intended use case is a cheap execution node for tool calls and high-volume mechanical steps, while larger models handle the reasoning.
- The post notes it is API-only for now and not open weights; the previous Ling-2.6-flash had MIT licensing.
- The broader question raised is whether the industry is converging on a split between small fast executor models and larger planners.
Related event: Ling-3.0-flash Released: Targeting Agents with Long Free Access(10 posts)→
More from Infra
- AI demand is pushing DRAM prices up 171.8% and relief may not come until 2028 — 量子位 · 2026-07-25
- Reproducible backends could make GPU and CPU profiling portable, but the business model is unclear — omojumiller · 2026-07-25
- A practical document-parser playbook says layout and scan quality matter more than rankings — emmettvance · 2026-07-25
- NURL nears v1.0 with a single-binary language built for LLM workflows — AdhesivenessHappy873 · 2026-07-25
- OpenAI status page shows elevated errors across ChatGPT, APIs and Codex — Primary_Tea3095 · 2026-07-25
- Anthropic says it has supply deals with Samsung Electronics and SK hynix — dejavucoder · 2026-07-25