Microsoft’s Mage-VL cuts visual tokens by 75% in streaming multimodal tasks
microsoft · hf · 2026-07-29
Microsoft presents Mage-VL, a codec-native streaming multimodal foundation model for real-time understanding and interaction. The model uses a custom tokenizer, Mage-ViT, to selectively encode dynamic, entropy-rich regions instead of uniformly sampling frames.
- Efficiency: visual token usage drops by 75%+ while preserving spatiotemporal context.
- Training scale: about 560M unlabeled images and 100M unlabeled video frames.
- Architecture: a lightweight System 1 event gate plus a causal System 2 decoder.
- Results: Mage-VL-4B matches Qwen3-VL-4B on static tasks, improves video and spatial reasoning, and delivers up to 3.5x wall-clock speedup.
- Extra: the paper also reports seven empirical findings on data efficiency and multimodal training pipelines.
Related event: Microsoft unveils Mage-VL, a 4B streaming multimodal model(5 posts)→
More from Multimodal
- SenseNova Infographic-V3 Demo Released: Local Editing & Global Style Transformation — Secret_Yak2496 · 2026-07-30
- Pro Tip: Steer AI Video Angles Precisely Using Simple Camera Placement Diagrams — round · 2026-07-30
- Wonder: Real-Time Camera-Controllable World Model at 16 FPS — qixing_huang · 2026-07-30
- Bottlenecks in Long Video Outpainting: Splicing and Color Consistency — Alex-edits123 · 2026-07-30
- Contour: Open-Source Tool Turns 2D Maps into 3D Terrain with Gemini Voice Guide — tom_doerr · 2026-07-30
- Dev Uses AI to Build ComfyUI Node for Automatic LoRA Trigger Word Replacement — TrueRedditMartyr · 2026-07-30