Microsoft’s Mage-VL cuts visual tokens by 75% in streaming multimodal tasks

microsoft · hf · 2026-07-29

Microsoft presents Mage-VL, a codec-native streaming multimodal foundation model for real-time understanding and interaction. The model uses a custom tokenizer, Mage-ViT, to selectively encode dynamic, entropy-rich regions instead of uniformly sampling frames.

Related event: Microsoft unveils Mage-VL, a 4B streaming multimodal model(5 posts)→

Original post →

More from Multimodal

Multimodal channel →