inclusionAI Releases Ling-3.0-tiny: An 8B Parameter MoE Model

-Cubie- · reddit · 2026-08-11

inclusionAI has open-sourced Ling-3.0-tiny, a much smaller version of Ling-3.0-flash, on Hugging Face. The model features 8B total parameters with only 1.3B active parameters.

Performance-wise, it falls between 4B and 12B Qwen and Gemma models. Thanks to the tiny active parameter count, it is expected to deliver massive tokens/sec inference speeds on most systems, highlighting the potential of tiny MoE architectures.

Related event: Ant's Ling Team Open-Sources 8B MoE Model Ling-3.0-tiny(2 posts)→

Original post →

More from Models

Models channel →