Cohere's Tiny Aya Vision: sub-4B multilingual VLM covering 70+ languages
Cohere_Labs · x · 2026-08-24
Cohere Labs Community researchers introduce Tiny Aya Vision, an open-weight multilingual vision-language model under 4B parameters supporting 70+ languages — the first of its kind at this scale. It extends Tiny Aya (3.35B, 70+ languages) with lightweight visual capabilities via parameter-efficient fusion.
The core question: can a 3B multilingual model gain effective visual grounding without sacrificing multilingual text performance or on-device deployability? No existing sub-4B model combines vision with 70+ language support (vs. Qwen3-VL-2B, Gemma 3-1B, SmolVLM, etc.).
Method: a frozen encoder, a small connector, and LoRA, then merging weights back to restore text performance. At 8B, this merging lifted multilingual vision win-rate against Pangea-7B from 58.1% to 70.0% on AyaVisionBench; the open question is whether the gain holds at 3.35B. Code is open-sourced; results reflect an in-progress v0 model.
More from Multimodal
- Gemini 3.7 Flash Natively Supports Video Transcription and Object/Action Recognition — DynamicWebPaige · 2026-08-25
- Call for Best Local VLMs - August 2026 — rm-rf-rm · 2026-08-25
- Flux1 images to WAN2.2 clips: how to stitch them into 15+ second videos — wreck_of_u · 2026-08-24
- ComfyUI Blue Run Button Executes Whole Workflow Instead of Branch — ImaginaryIncident481 · 2026-08-24
- MiniMax M3 transforms logos into premium brand films with minimal effort — petewoodbridge · 2026-08-24
- InfinityEdit enables infinite video editing with a lightweight adapter — _akhaliq · 2026-08-24