Alibaba Releases Swift-Image: A Compact 6B Unified Text-to-Image and Editing Model
HaktanSuren · x · 2026-08-25
Alibaba's team releases Swift-Image, a compact unified model covering text-to-image generation, single-image editing, and multi-image editing.
- Built on an efficient 6B single-stream DiT with a progressive training pipeline moving from broad semantic coverage toward higher resolution and stronger visual quality
- Post-training combines parallel expert RL + multi-teacher on-policy distillation to reduce interference among heterogeneous objectives
- A Prompt Enhancer translates user requests into generator-aligned visual specs, decoupling high-level reasoning from pixel rendering
- Structural pruning and few-step distillation yield a 3B variant (nearly lossless) and accelerated versions
- With only 6B parameters and 243K GPU training hours, it achieves leading aggregate performance among evaluated open-source models
Related event: Alibaba Releases Swift-Image 6B Unified Image Generation and Editing Model(2 posts)→
More from Multimodal
- Using NAIv5 in-context learning to generate images in personal style — Birchlabs · 2026-08-25
- Seedance 2.5 still misspells text in videos despite prompt emphasis — AthleteArtistic3121 · 2026-08-25
- NoSpoon Studios demos AI-generated music video in private alpha — Kyrannio · 2026-08-25
- Studio 1939 LoRA released for Minimax H3 model — Affectionate-Map1163 · 2026-08-25
- Studio 1939 LoRA updated for Minimax H3 with better results — Affectionate-Map1163 · 2026-08-25
- User shares Comfy Minimax H3 workflow for one-shot fight scene generation — DifficultyNo4237 · 2026-08-25