Ling 3.0-flash-VL lands in llama.cpp: 124B MoE with only 5.5B active params
jacek2023 · reddit · 2026-09-24
- inclusionAI's Ling-3.0-flash-VL has a merged PR in llama.cpp (#29151), inheriting Ling-3.0-flash's language, reasoning, and long-context abilities while adding native image and video understanding.
- Key specs: 124B total parameters with only 5.5B activated per token, and context windows up to 256K tokens.
- Architecture highlights:
- A ViT visual encoder with a two-layer MLP projector aligning visual and text features;
- VideoRoPE encodes spatial position and temporal order, enabling event localization, long-video QA, and clip editing;
- A 42-layer hybrid backbone alternating KDA and Gated MLA layers at 5:1 for efficient long-context processing across text, images, video, and long agent histories;
- Sparse MoE balances strong multimodal capability with inference efficiency.
More from Infra
- Kaivid Labs Migrates 400K+ Vectors From Milvus to Qdrant, Benchmarks Both — qdrant_engine · 2026-09-24
- Brookings: AI data center investment to hit $10.3 trillion between 2025 and 2032 — rohanpaul_ai · 2026-09-24
- Hyperscalers Need 2.7x Productivity Gains to Justify $1.1T AI Infrastructure Spend by 2030 — Post-reality · 2026-09-24
- HSIR: a Modal-like runtime for scientific AI that keeps data in your VPC, 55.8s on DAX benchmark — retr0jirachi · 2026-09-24
- Cloudflare reclaims 100TB of RAM with math and Rust in Pingora — bibryam · 2026-09-24
- Intel Arc B580 runs INT8 ConvRot acceleration, Flux 2 Klein 9B in 3.9s — Valuable-Subject-274 · 2026-09-24