Ant's open-source Ling-3.0-flash-VL edges out GPT-5.4 on Image-to-WebDevArena

智东西 · wechat · 2026-09-09

Ant Group open-sourced its first natively multimodal model, Ling-3.0-flash-VL: a 124B-parameter MoE (5.5B active) with native image/text/video input, 256K context, and BF16/FP8 weights on HuggingFace and ModelScope.

Overseas devs already recreated a playable Flappy Bird page from a screenshot, calling it "a pretty solid release."

Related event: Ant Group Open-Sources Ling-3.0-flash-VL Multimodal Model(4 posts)→

Original post →

More from coding & agent

coding & agent channel →