Microsoft’s Mage-Flow packs 4B image generation and editing into interactive speeds
microsoft · hf · 2026-07-22
Microsoft introduces Mage-Flow, a compact 4B image generation and editing stack built for native-resolution training and deployment.
What it includes
- Mage-VAE: a lightweight latent tokenizer that uses one-step diffusion-style encoding/decoding and cuts tokenization cost by more than an order of magnitude.
- Native-Resolution Multimodal Diffusion Transformer: trained with rectified flow matching for flexible-resolution generation.
- Stack-level CUDA kernel fusion and native-resolution packing raise end-to-end training throughput by about 2.5×.
Reported performance
- The family includes Base, RL-aligned, and Turbo variants for both generation and editing.
- Diffusion-NFT improves prompt following, text rendering, aesthetics, and editing fidelity.
- At 1024² on a single NVIDIA A100, Mage-Flow-Turbo generates an image in 0.59s and Mage-Flow-Edit-Turbo edits one in 1.02s.
Microsoft says the result is a practical high-resolution generation/editing system that stays small enough for interactive use.
Related event: Microsoft Asia Open-Sources Mage-Flow for Image Generation and Editing(3 posts)→
More from Multimodal
- SIGGRAPH 2026 workshop will cover generative AI across 3D, simulation and animation — qixing_huang · 2026-07-22
- Musk backs Grok Imagine as AI video starts to feel culturally useful — elonmusk · 2026-07-22
- Krea 2 Identity Edit shows stronger identity preservation in image edits — Fishmongr · 2026-07-22
- Fable 5’s broad safeguards flag routine coding and biology work, then switch to Opus 4.8 — sumitdotml · 2026-07-22
- A 3-year throwback shows how rough AI video used to be — Confident_Salt_8108 · 2026-07-22
- Open-source skill turns Chinese stories into hand-drawn diary videos — dotey · 2026-07-22