A New Benchmark for Evaluating VLM ASCII Art
East-Muffin-6472 · reddit · 2026-07-19
This is a new benchmark named ASCIITermDraw-Bench, designed to test whether VLMs can genuinely generate and edit ASCII diagrams as instructed.
The author's motivation: while many models can "describe" what an architecture or topology diagram looks like, accurately formatting that content into ASCII text requires a different, harder capability. The benchmark features 80 tasks across four categories:
- Basic boxes and layouts
- Network topologies
- Software architecture diagrams
- Image-conditioned ASCII editing (modifying a diagram while preserving unrequested parts)
The evaluation methodology is comprehensive:
- Structural Score: Checks if necessary labels, edges, entities, and relationships are correct.
- Semantic Score: Evaluated by an LLM judge, repeating each task 5 times to minimize variance.
- Final results aggregate the 80 tasks, providing a 95% confidence interval.
In the current leaderboard, Gemma-4-31B-IT leads with 73.8%, followed by Qwen3.7-Plus (70.2%), Kimi-K2.6 (61.8%), and MiniMax-M3 (59.5%). The project has open-sourced 12 sample tasks and the complete methodology, with reproducible content available on Hugging Face.
Related event: ASCIITermDraw-Bench Tests VLMs on ASCII Drawing(2 posts)→
More from Multimodal
- HeyGen adds a media-sourcing skill for coding agents with 75k images and 10k tracks — HeyGen · 2026-07-22
- Reddit user seeks ComfyUI NSFW text-to-image and image-to-video workflows under 20 GB VRAM — hobbyist2020 · 2026-07-22
- Krea 2 users recommend a two-pass Clownshark sampler setup for sharper image details — listopalafoto · 2026-07-22
- Gemini Omni Flash turns a boat cabin into a cave in Flow by Google — chrisfirst · 2026-07-22
- A simple workflow to turn a photo into an image prompt using Gemini, Grok, or GPT Image — harshitagu72595 · 2026-07-22
- A Reddit user proposes a consistency LoRA to keep anime and game scenes visually stable — ThirdWorldBoy21 · 2026-07-22