Microsoft’s first text-to-image model scores 49% in a 192-prompt benchmark but lags on realism
dh7net · reddit · 2026-07-27
A Reddit user benchmarked Microsoft’s first text-to-image model, Mage-Flow-Turbo, and compared it with other 4B open models.
What the model is
- 4B parameters, MIT licensed
- Native-resolution generation from 512 to 2048 px at any aspect ratio
- A 4-step distilled turbo model
- On the author’s DGX Spark, it can generate a 1024×1024 image in about 4.6 seconds
How it performed
The user evaluated it on 192 prompts across six categories — text, spatial reasoning, human realism, truthfulness, studio/product, and graphic design — with image-by-image judging by Gemini 3.1 Pro.
- It slightly beats Flux 2 Klein 4B on capability score: 49% vs 48%
- But Klein wins on aesthetics, so its overall score is higher: about 51 vs 47
- Another 4B open model, Bonsai Ternary 4B, comes in just behind both
Where Mage-Flow-Turbo shines
- Professional/studio/product shots: 85%
- Text rendering: 67%
- Speed
Where it struggles
- Human realism: 29%
- Truthfulness / world knowledge: 37%
- Spatial reasoning: 49%
The author concludes that it may mainly be useful for people constrained to 4B due to VRAM limits.
More from Models
- A user says Claude’s $100 monthly plan hits usage limits in two days — talkaboutdesign · 2026-07-27
- Meta is reportedly preparing a harness and new open-source models — garrytan · 2026-07-27
- Sonnet completely melts down on a 192×191×190×…×1 arithmetic prompt — PipeTasty7582 · 2026-07-27
- User says 5.6 Sol High beats Fable on research and costs €23 a month — PressPlayPlease7 · 2026-07-27
- llama.cpp merges GLM-5.2-Vision support for local multimodal inference — QuixiAI · 2026-07-27
- Elon Musk calls Grok 4.5 a solid workhorse after Tim Sweeney’s praise — elonmusk · 2026-07-27