inclusionAI open-sources Ling-3.0-flash-VL: 124B MoE with 5.5B active and 1M context
jacek2023 · reddit · 2026-09-08
inclusionAI released Ling-3.0-flash-VL on Hugging Face:
- Sparse MoE with 124B total parameters, only 5.5B activated per token, supporting up to 1M-token context;
- Inherits Ling-3.0-flash's language, reasoning, and long-context strengths, adding native image and video understanding for real-world reasoning and agentic workflows;
- Architecture: ViT encoder + two-layer MLP projector; VideoRoPE encodes spatial position and temporal order for event localization, long-video QA, and clip editing; a 42-layer hybrid backbone alternates KDA and Gated MLA layers at a 5:1 ratio for efficient multimodal long-context processing.
More from Models
- Qwen quietly releases Drive-1.0-4B, a 4B driving model finetuned from Qwen3.5 — FullstackSensei · 2026-09-09
- Inception ships Mercury 2.5: diffusion LLM with 40% intelligence jump at 1,100 tokens/sec — timshi_ai · 2026-09-09
- OpenAI claims agent-produced solution to 90-year-old Navier-Stokes Millennium Problem — thesaraharminta · 2026-09-09
- OpenAI says new model solved Navier–Stokes in 88 hours with 10,000 coordinating agents — legit_api · 2026-09-09
- Skeptic mocks OpenAI's containment claim: couldn't even manage 1,000 instances, now 10,000 — scaling01 · 2026-09-09
- Gary Marcus mocks GPT-6 Astra video: OpenAI's 'human-level intelligence' claim is a lie — GaryMarcus · 2026-09-09