Ling-3.0-tiny Released: 1.3B Active Params Beats GPT-OSS 120B
solyarisoftware · x · 2026-08-11
Ling-3.0-tiny is an open-weight model designed for reasoning, tool calling, and agentic workflows. It uses a 128-expert sparse MoE architecture with 7.9B total parameters but only 1.3B active parameters per token.
- Performance: Scores 73.4 on GPQA Diamond and 71.0 on IMO-AnswerBench. It reportedly beats OpenAI's GPT-OSS 120B (high) on Artificial Analysis despite its small size.
- Local Deployment: Available in BF16, FP8, and INT4. The INT4 version requires only 5.8GB memory, while FP8 peaks at 8.34 GiB for 8K context.
- Inference Speed: Achieves 100-105 tok/s on DGX Spark and 86-90 tok/s on M4 Pro MacBook.
- Runtime Support: Currently supports SGLang, vLLM, and MLX on Apple Silicon. Support for llama.cpp and LM Studio is pending.
Related event: Ant's inclusionAI Open-Sources Ling-3.0-tiny MoE Model(3 posts)→
More from Infra
- Under 10% of Enterprises Scale AI; Compute Shortage to Persist — BenBajarin · 2026-08-11
- HKU Team Bypasses Von Neumann Bottleneck with Ultra-low Power 2D Material AI Chip — YiMaTweets · 2026-08-11
- Blockstream launches atomic swaps for Bitcoin and Lightning, citing AI-assisted attacks — RSync25 · 2026-08-11
- Local AI Video on RTX 3060 Ti: Generate High-Quality Clips in 20 Minutes with 8GB VRAM — cocktailpeanut · 2026-08-11
- Running 35B Models on a Single RTX 5080: Migrating from llama.cpp to vLLM — McFlurriez · 2026-08-11
- SlimServe Open-Sourced: Runs DeepSeek at 1k tok/s on 4x A100s — QuixiAI · 2026-08-11