Ant Group Open-Sources Ling-3.0-tiny: 1.3B Active Params Beats 31B Models
智东西 · wechat · 2026-08-11
Ant Group's Bailian team has open-sourced Ling-3.0-tiny, a lightweight hybrid reasoning MoE model designed for efficient local deployment. It has 7.9B total parameters but only activates 1.3B during inference.
- Architecture: Uses a 3:1 alternating stacking of KDA (Kimi Delta Attention) and MLA (DeepSeek's Multi-head Latent Attention), combined with 128 sparse experts to balance long-context capabilities and compute costs.
- Performance: Scores 25 on the ArtificialAnalysis index, closely trailing Gemma-4-26B, and surpasses Gemma-4-31B in agentic tasks.
- Edge Deployment: Offers BF16, FP8, and INT4 versions. Verified on DGX Spark, MacBook, and Mac mini. The FP8 version achieves 86-90 tokens/s on M4 Pro MacBooks with first-token latency under 100ms.
More from coding & agent
- A YC founder's early sales guide for vibe coders: LinkedIn caps you at 200 requests/week — namanyayg · 2026-08-26
- Distinguishing fact from hallucination in MCP agent audits — saas-wizard · 2026-08-26
- Don't let LLMs decide who can write to main — saas-wizard · 2026-08-26
- Open-source tool converts YouTube videos into structured Obsidian Markdown notes — tom_doerr · 2026-08-26
- Moving the verdict outside the model for explainability — Jay299792458 · 2026-08-26
- Ox Alpha processes 11.6T tokens in three days — rohanpaul_ai · 2026-08-26