Ling-3.0-flash Open-Sourced: 127B MoE with Official FP8 Requiring Only 128GB
derspenti · reddit · 2026-08-04
inclusionAI has released the weights for Ling-3.0-flash on Hugging Face under the MIT license.
- Architecture & Parameters: Total parameters 127.5B with 5.1B active. Uses BailingMoeV3 architecture with 512 experts (8 active per token), offering finer granularity than most peers.
- Thinking Mode: Can be toggled per-request inside the chat template, defaulting to on, eliminating the need for a separate SKU.
- Local Deployment: BF16 version is 255GB; the official FP8 version is 128GB, making it a direct download for large unified-memory boxes or multi-GPU rigs. Community is currently discussing llama.cpp compatibility.
Related event: Ling-3.0-flash Hybrid MoE Model Goes Open Source(2 posts)→
More from Infra
- Tested: MiniMax H3 Runs Locally on 64GB MacBook Pro via Phosphene — cocktailpeanut · 2026-08-05
- NVIDIA Joins NSF Regional AI Hubs to Expand Computing Access Nationwide — nordicinst · 2026-08-05
- Agentic RL Bottlenecked by Inference: SkyPilot Halves Training Time — skypilot_org · 2026-08-05
- Hardware Architecture Debate: Why Vertical Power Delivery Over Vertical Optical IO? — jwt0625 · 2026-08-04
- CoreWeave Announces Fully Connected 2026: Fei-Fei Li & NVIDIA to Keynote — wandb · 2026-08-04
- Agentic AI Triggers a Storage Shock: Enterprise Data Becomes the New Bottleneck — BenBajarin · 2026-08-04