Microsoft’s first text-to-image model scores 49% in a 192-prompt benchmark but lags on realism

dh7net · reddit · 2026-07-27

A Reddit user benchmarked Microsoft’s first text-to-image model, Mage-Flow-Turbo, and compared it with other 4B open models.

What the model is

How it performed

The user evaluated it on 192 prompts across six categories — text, spatial reasoning, human realism, truthfulness, studio/product, and graphic design — with image-by-image judging by Gemini 3.1 Pro.

Where Mage-Flow-Turbo shines

Where it struggles

The author concludes that it may mainly be useful for people constrained to 4B due to VRAM limits.

Original post →

More from Models

Models channel →