Microsoft Unveils Mage-VL: Codec-Native Streaming VLM with 3.5x Inference Speedup
pmttyji · reddit · 2026-07-29
Microsoft introduced Mage-VL, a 4B parameter codec-native streaming multimodal foundation model designed for efficient image and video understanding.
- Core Innovation: The visual encoder is trained entirely from scratch. Borrowing the I/P frame mechanism from video codecs, it cuts visual tokens by over 75% by keeping anchor frames and only extracting patches from regions with actual motion.
- Inference Speedup: Achieves up to 3.5x wall-clock inference speedup over uniform frame sampling at matched accuracy, featuring native-resolution scaling.
- Dual-Process Design: Implements a System 1 (a lightweight cognition gate for low-latency monitoring) and System 2 (the full VLM for complex events) within a single model to enable proactive streaming.
- Performance: With a fixed Qwen3-4B LLM backbone, it significantly outperforms the similarly-sized Qwen3-VL-4B across video understanding and spatial intelligence benchmarks.
Related event: Microsoft Launches Mage-VL Codec-Native Streaming Multimodal Model(3 posts)→
More from Models
- GPT-5.6 Sol Ultra finds a critical bug, then refuses to show it — haltakov · 2026-07-29
- Kimi K3 tops a benchmark chart in a repost claiming it beats Anthropic models — JarnoDuursma · 2026-07-29
- User Reports Grok's Generation Capabilities Have Gotten 'Real Cracked' — djcows · 2026-07-29
- User asks Anthropic not to deprecate Opus 4.6 until the model is fixed — oyacaro · 2026-07-29
- Hidden Trick: Manually Invoke Older Opus Models in Claude Code — voooooogel · 2026-07-29
- Claude Sonnet 3 is set to sunset on July 30 — repligate · 2026-07-29