Microsoft’s Mage-VL cuts visual tokens by 75% in streaming multimodal tasks
microsoft · hf · 2026-07-29
Microsoft presents Mage-VL, a codec-native streaming multimodal foundation model for real-time understanding and interaction. The model uses a custom tokenizer, Mage-ViT, to selectively encode dynamic, entropy-rich regions instead of uniformly sampling frames.
- Efficiency: visual token usage drops by 75%+ while preserving spatiotemporal context.
- Training scale: about 560M unlabeled images and 100M unlabeled video frames.
- Architecture: a lightweight System 1 event gate plus a causal System 2 decoder.
- Results: Mage-VL-4B matches Qwen3-VL-4B on static tasks, improves video and spatial reasoning, and delivers up to 3.5x wall-clock speedup.
- Extra: the paper also reports seven empirical findings on data efficiency and multimodal training pipelines.
Related event: Microsoft unveils Mage-VL, a 4B streaming multimodal model(5 posts)→
More from Multimodal
- Opus 5.5 turns a single image into a Three.js game menu in one simple prompt — majidmanzarpour · 2026-09-23
- PixVerse's R2 world model goes hands-on: endless exploration, but compute limits cap play time — Xianbao_QIAN · 2026-09-23
- Local image generation tests: Qwen base model combined with a Flux refiner — freshstart2027 · 2026-09-23
- Reddit users find Qwen 2.1 a major disappointment for text-to-image — -becausereasons- · 2026-09-23
- Given a 3-hour budget and a one-line prompt, Opus 5.5 produced a full Kowloon horror short itself — rainbird · 2026-09-23
- Image-gen prompt sharer techhalla posts 'safer highways' generations with prompts — techhalla · 2026-09-23