Text-to-Image benchmark with 9k images tests 52 models

dh7net · reddit · 2026-08-27

A new Text-to-Image benchmark features 192 prompts designed to be difficult for T2I models (text rendering, spatial reasoning, human realism, negations), judged by a VLM against ground truth.

It covers 52 models with over 9,000 generated and analyzed images. Unlike many leaderboards, the full dataset (including prompts and results) is published on Hugging Face, with a public gallery for visual inspection.

Links:

Original post →

More from Multimodal

Multimodal channel →