Ant Group Open-Sources Ling-3.0-flash-VL Multimodal Model
Ant Group's inclusionAI open-sourced Ling-3.0-flash-VL, a 124B sparse MoE vision-language model activating only 5.5B parameters per token, supporting image/video/text and long context, with BF16 and FP8 weights released and visual coding scores surpassing GPT-5.4.
2026-09-08 ~ 2026-09-09 · 4 related posts
- inclusionAI open-sources Ling-3.0-flash-VL: 124B MoE with 5.5B active and 1M context — jacek2023 · 2026-09-08
- inclusionAI open-sources vision model Ling-3.0-flash-VL with BF16 and FP8 weights — FellMentKE · 2026-09-09
- Ant's Ling-3.0-flash-VL Open-Sourced: A Vision Model That Uses Tools — FellMentKE · 2026-09-09
- Ant's open-source Ling-3.0-flash-VL edges out GPT-5.4 on Image-to-WebDevArena — 智东西 · 2026-09-09