Microsoft ships Mage-VL 4B on Hugging Face as a streaming video-language model
multimodalart · x · 2026-07-27
Microsoft has released Mage-VL 4B on Hugging Face, describing it as an efficient codec-native streaming VLM.
The model can describe live events as they happen and can be prompted to focus on a specific moment, such as the instant a train arrives or when a goal is scored. A Spaces demo is available for testing.
More from Multimodal
- AI-Generated Cat Adventure Videos Go Viral with Over 20M Views — aziz4ai · 2026-07-28
- Hugging Face open-sources a local speech-to-speech voice-agent stack — huggingface · 2026-07-28
- AI Rebuilds Homer's Odyssey into a 135-Minute Feature Film — heypearlai · 2026-07-28
- An AI remake of The Odyssey trailer scenes is weird enough to be worth sharing — eyishazyer · 2026-07-28
- Nvidia puts an open vision-language-action model on Hugging Face — theteknosaur · 2026-07-28
- A Grok user asks to turn Hodor into Shrek in a classic AI meme prompt — heypearlai · 2026-07-28