A New Benchmark for Evaluating VLM ASCII Art
East-Muffin-6472 · reddit · 2026-07-19
This is a new benchmark named ASCIITermDraw-Bench, designed to test whether VLMs can genuinely generate and edit ASCII diagrams as instructed.
The author's motivation: while many models can "describe" what an architecture or topology diagram looks like, accurately formatting that content into ASCII text requires a different, harder capability. The benchmark features 80 tasks across four categories:
- Basic boxes and layouts
- Network topologies
- Software architecture diagrams
- Image-conditioned ASCII editing (modifying a diagram while preserving unrequested parts)
The evaluation methodology is comprehensive:
- Structural Score: Checks if necessary labels, edges, entities, and relationships are correct.
- Semantic Score: Evaluated by an LLM judge, repeating each task 5 times to minimize variance.
- Final results aggregate the 80 tasks, providing a 95% confidence interval.
In the current leaderboard, Gemma-4-31B-IT leads with 73.8%, followed by Qwen3.7-Plus (70.2%), Kimi-K2.6 (61.8%), and MiniMax-M3 (59.5%). The project has open-sourced 12 sample tasks and the complete methodology, with reproducible content available on Hugging Face.
Related event: ASCIITermDraw-Bench Tests VLMs on ASCII Drawing(2 posts)→
More from Multimodal
- FLUX.2 Klein Drifts Hard on Character Expressions While Free Gemini Holds Likeness — wacomlover · 2026-09-11
- Tencent Hunyuan releases AuK code and weights on GitHub with ComfyUI and fine-tuning support — aigclink · 2026-09-11
- Tencent open-sources AuK, a unified 1.5B speech generation and editing model — aigclink · 2026-09-11
- Creator turns Bahamut vs Tiamat rivalry into an AI cinematic battle with Midjourney, GPT Image 2 and Seedance — azed_ai · 2026-09-11
- invideo launches AI agent-powered editor to automate repetitive editing tasks — azed_ai · 2026-09-11
- fable 5.1 recreates The Starry Night with 256,157 JavaScript brush strokes — cedric_chee · 2026-09-11