Ant's Ling releases Ling-3.0-flash-VL, a free multimodal model trained jointly on text, images, video
nikola_mr64990 · x · 2026-09-21
Ant Group's Ling team released Ling-3.0-flash-VL, its first multimodal model, currently free to use via OpenRouter.
- Not a text model with a bolted-on vision head: images, text, and video share a single reasoning chain from training onward
- Aimed at understanding screenshots, layouts, and visual context for everyday development and creative workflows
More from Multimodal
- Qwen-Image 2.1 LoRA training plagued by anatomy failures, community hunts for fixes — ResidentFrame4195 · 2026-09-21
- Seedance 2.5 reskins an existing animation shot-for-shot, sparking workflow debate — jjvincent · 2026-09-21
- Minimax H3 generates a startlingly realistic Big Bang Theory Penny clip — blackdatafilms · 2026-09-21
- One screenshot, one prompt: Ling-3.0-flash-VL designs a foldable device concept where hinge angle is the UI — nikola_mr64990 · 2026-09-21
- Rodin-generated 3D instruments + GPT-6 build a playable interactive instrument website in 4 steps — CurieuxExplorer · 2026-09-21
- ComfyUI-Kompressor v1.0 ships INT8/W4A8 ConvRot quantization for FLUX.2, Wan 2.2, and more — a3tinu · 2026-09-21