UltraText Bench: bilingual dense text rendering eval, Qwen-Image drops from 86.50 to 42.86 at L3
Westlake-University · hf · 2026-10-08
Westlake University's LINs lab introduces UltraText Bench, a bilingual benchmark for dense visual text rendering in prompt-only image generation, testing sustained performance across demanding scenes as short-string rendering improves.
Setup:
- 432 human-reviewed prompts across 24 real-world scene categories and three difficulty levels, split evenly between English and Chinese;
- Each prompt supplies exact strings for 4-12 text regions plus structured references for content, placement, and visual attributes;
- Q-Judger VLM scores text fidelity, clarity, spatial quality, and scene quality, validated by a 10-person human evaluation.
Findings: across 24 model configurations, dimensions diverge—Z-Image-Turbo gains 3.81 clarity points over Z-Image-Base while losing 14.76 fidelity points; performance degrades with workload, as Qwen-Image-2512's English composite falls from 86.50 at L1 to 42.86 at L3. Repository on GitHub.
More from Multimodal
- KAIST's Tetris3D Reconstructs 3D Scenes With Physically Coherent Objects, Plus 1.2M-Scene ComOb Dataset — kaist-ai · 2026-10-08
- Suno Praised as Good Enough to Eventually Overtake Spotify — petergyang · 2026-10-08
- 3.5M-param anime upscaler applies Looped-DiT recurrence, gains +0.6dB in 4 loops — NobodySnJake · 2026-10-08
- HKU and VAST's Mira-Scene fixes object placement in single-image 3D scene reconstruction — jiqizhixin · 2026-10-08
- Musk Hypes Grok Bot After User Generates Full Explainer Video From a 15-Second Prompt — elonmusk · 2026-10-08
- AI pop singer Claudia open-sourced: downloadable soul via MCP server, skills and character sheets — Promptmethus · 2026-10-08