Ant's Ling-3.0-flash-VL Open-Sourced: A Vision Model That Uses Tools
FellMentKE · x · 2026-09-09
Ant's AntLingAGI open-sourced Ling-3.0-flash-VL in BF16 and FP8 (FP4/INT4 coming).
- Goes beyond recognition: understands images, video, docs and UIs
- Reasons, searches and verifies from visual cues; can call tools, check results and deliver outcomes
- Demos show phone-camera use: object recognition, scene narration, foreign-sign translation for travel and accessibility
- Commenters highlight no-code visual editing workflows as a practical use.
More from Multimodal
- Meta quietly shipped Muse Spark 1.3 with no blog post, alongside Astra and Fable — aparnadhinak · 2026-09-09
- ChatGPT images recreate iPhone 18 Pro Max Cherry product shots before Apple's launch — techhalla · 2026-09-09
- Z.ai demos 3D character exploration powered by an LLM (jokingly dubbed GPT-2.5) — azed_ai · 2026-09-09
- Reddit user demos Wan 3.0 image-to-video, chaining references to keep character consistent — NatalieCrypto · 2026-09-09
- One prompt, one playable world: Astra + Thrixel build a Three.js alien marketplace — RanaHanocka · 2026-09-09
- MiniMax H3 image-to-video auto-stretches output at 1344x768 but works fine at 1056x608 — Guyserbun007 · 2026-09-09