Microsoft’s Mage-Flow packs 4B image generation and editing into interactive speeds
microsoft · hf · 2026-07-22
Microsoft introduces Mage-Flow, a compact 4B image generation and editing stack built for native-resolution training and deployment.
What it includes
- Mage-VAE: a lightweight latent tokenizer that uses one-step diffusion-style encoding/decoding and cuts tokenization cost by more than an order of magnitude.
- Native-Resolution Multimodal Diffusion Transformer: trained with rectified flow matching for flexible-resolution generation.
- Stack-level CUDA kernel fusion and native-resolution packing raise end-to-end training throughput by about 2.5×.
Reported performance
- The family includes Base, RL-aligned, and Turbo variants for both generation and editing.
- Diffusion-NFT improves prompt following, text rendering, aesthetics, and editing fidelity.
- At 1024² on a single NVIDIA A100, Mage-Flow-Turbo generates an image in 0.59s and Mage-Flow-Edit-Turbo edits one in 1.02s.
Microsoft says the result is a practical high-resolution generation/editing system that stays small enough for interactive use.
Related event: Microsoft Asia Open-Sources 4B Image Model Mage-Flow(12 posts)→
More from Multimodal
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- MiniMax H3 MAX nails cooking anime clips: 15-second curry demo with prompts shared — Hailuo_AI · 2026-09-11
- Using a finisher move on one mosquito with MiniMax H3 MAX — the bug survives — Hailuo_AI · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Lumara AI Film Festival Comes to NYC Oct 26, Top AI Filmmakers to Compete — 0xAllen_ · 2026-09-11
- Pterodactyl Detective: An AI-Generated Proof-of-Concept Trailer — PterodactylDetective · 2026-09-11