Ant's open-source Ling-3.0-flash-VL edges out GPT-5.4 on Image-to-WebDevArena
智东西 · wechat · 2026-09-09
Ant Group open-sourced its first natively multimodal model, Ling-3.0-flash-VL: a 124B-parameter MoE (5.5B active) with native image/text/video input, 256K context, and BF16/FP8 weights on HuggingFace and ModelScope.
- Visual feedback loop: the model re-reads page/environment state after acting and iterates—screenshot → code → browser render → compare → refine; the same mechanism powers GUI agent workflows across apps.
- Benchmarks: 42 on Artificial Analysis index (+4 vs text-only); 1441 on Image-to-WebDevArena vs GPT-5.4's 1440; OmniDocBench 91.35, WebVoyager 90.83—though still behind frontier closed models like Claude Fable 5.1 (66).
- Use cases: sports highlight auto-editing, real-time camera Q&A in the Ling Guang app, structured extraction across multi-page medical reports (explicitly not a doctor replacement).
Overseas devs already recreated a playable Flappy Bird page from a screenshot, calling it "a pretty solid release."
Related event: Ant Group Open-Sources Ling-3.0-flash-VL Multimodal Model(4 posts)→
More from coding & agent
- Fudan's Feedback-Enriched Environments Bootstrap Self-Evolving Agents in Long-Horizon Tasks — FudanUniversity · 2026-09-09
- Even the local barber is using Claude Code now — garrytan · 2026-09-09
- Meta details Muse agent safety: sandboxed harness, Sentinel gatekeeper, user-held encryption — unixterminal · 2026-09-09
- GitHub Copilot CLI bug: Mission Control dashboard links 404 as sessions live under /agents/tasks path — dai · 2026-09-09
- SREGym benchmark launches: GPT-5.6 leads agents at fixing real production SRE failures — tianyin_xu · 2026-09-09
- Security blogs warn of AI doom but no lab has published a guide to audit your own security — RhysSullivan · 2026-09-09