Microsoft ships Mage-VL 4B on Hugging Face as a streaming video-language model
multimodalart · x · 2026-07-27
Microsoft has released Mage-VL 4B on Hugging Face, describing it as an efficient codec-native streaming VLM.
The model can describe live events as they happen and can be prompted to focus on a specific moment, such as the instant a train arrives or when a goal is scored. A Spaces demo is available for testing.
Related event: Microsoft unveils Mage-VL, a 4B streaming multimodal model(5 posts)→
More from Multimodal
- Third-party test: Claude Opus 5.5 renders finer 3D scenes but costs 13x more than GPT-6 Sol — testingcatalog · 2026-09-23
- ComfyUI trick: aux preprocessor + Qwen transfers poses across characters with one prompt — Acceptable-Work8202 · 2026-09-23
- Same portrait prompt across Midjourney V6.1, V7 and V8.2: do older models look better? — tisch_eins · 2026-09-23
- Testing AI character consistency across a 20-image travel sequence — SiennaVaire · 2026-09-23
- Midjourney v8.2 Faces: New Portrait Generation Samples Shared — azed_ai · 2026-09-23
- One-sentence prompt generates lifelike dog video, shown side-by-side with the real one — wgrathwohl · 2026-09-23