UltraText Bench: bilingual dense text rendering eval, Qwen-Image drops from 86.50 to 42.86 at L3

Westlake-University · hf · 2026-10-08

Westlake University's LINs lab introduces UltraText Bench, a bilingual benchmark for dense visual text rendering in prompt-only image generation, testing sustained performance across demanding scenes as short-string rendering improves.

Setup:

Findings: across 24 model configurations, dimensions diverge—Z-Image-Turbo gains 3.81 clarity points over Z-Image-Base while losing 14.76 fidelity points; performance degrades with workload, as Qwen-Image-2512's English composite falls from 86.50 at L1 to 42.86 at L3. Repository on GitHub.

Original post →

More from Multimodal

Multimodal channel →