Microsoft Unveils Mage-VL: Codec-Native Streaming VLM with 3.5x Inference Speedup
pmttyji · reddit · 2026-07-29
Microsoft introduced Mage-VL, a 4B parameter codec-native streaming multimodal foundation model designed for efficient image and video understanding.
- Core Innovation: The visual encoder is trained entirely from scratch. Borrowing the I/P frame mechanism from video codecs, it cuts visual tokens by over 75% by keeping anchor frames and only extracting patches from regions with actual motion.
- Inference Speedup: Achieves up to 3.5x wall-clock inference speedup over uniform frame sampling at matched accuracy, featuring native-resolution scaling.
- Dual-Process Design: Implements a System 1 (a lightweight cognition gate for low-latency monitoring) and System 2 (the full VLM for complex events) within a single model to enable proactive streaming.
- Performance: With a fixed Qwen3-4B LLM backbone, it significantly outperforms the similarly-sized Qwen3-VL-4B across video understanding and spatial intelligence benchmarks.
Related event: Microsoft unveils Mage-VL, a 4B streaming multimodal model(5 posts)→
More from Models
- Muse reportedly offers 4B tokens/week for ~$100/month, sparking industry price-disruption talk — NewYak4281 · 2026-09-23
- GPT-6 Sol and Luna appear in OpenAI docs, alongside guidance on reasoning effort — cedric_chee · 2026-09-23
- GPT-6 tested on LIBERO robot task: turns on stove, fails to grasp moka pot — YuXiang_IRVL · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23
- Claude Opus 5.5 costs $5.98 per task as price cuts offset ~80% token usage spike — ArtificialAnlys · 2026-09-23
- Claude Opus 5.5 benchmarked: intelligence 58, per-task cost spans 11x across five tiers — ArtificialAnlys · 2026-09-23