Microsoft releases Mage-VL on Hugging Face for streaming image and video understanding

_akhaliq · x · 2026-07-28

Microsoft has released Mage-VL on Hugging Face, describing it as a codec-native, proactive-streaming multimodal foundation model for image and video understanding.

The attached diagram shows an event-triggered commentary setup: a gate predictor decides whether to speak, then an LLM decoder generates commentary only when the gate opens. The model uses shared visual features from Mage-ViT and processes a continuous live-match stream to produce spoken commentary selectively.

In short, the release is aimed at streaming multimodal understanding rather than simple offline captioning.

Related event: Microsoft unveils Mage-VL, a 4B streaming multimodal model(5 posts)→

Original post →

More from Infra

Infra channel →