Cohere's Tiny Aya Vision: sub-4B multilingual VLM covering 70+ languages

Cohere_Labs · x · 2026-08-24

Cohere Labs Community researchers introduce Tiny Aya Vision, an open-weight multilingual vision-language model under 4B parameters supporting 70+ languages — the first of its kind at this scale. It extends Tiny Aya (3.35B, 70+ languages) with lightweight visual capabilities via parameter-efficient fusion.

The core question: can a 3B multilingual model gain effective visual grounding without sacrificing multilingual text performance or on-device deployability? No existing sub-4B model combines vision with 70+ language support (vs. Qwen3-VL-2B, Gemma 3-1B, SmolVLM, etc.).

Method: a frozen encoder, a small connector, and LoRA, then merging weights back to restore text performance. At 8B, this merging lifted multilingual vision win-rate against Pangea-7B from 58.1% to 70.0% on AyaVisionBench; the open question is whether the gain holds at 3.35B. Code is open-sourced; results reflect an in-progress v0 model.

Original post →

More from Multimodal

Multimodal channel →