New ASCIITermDraw benchmark says top VLMs still miss simple text diagrams
East-Muffin-6472 · reddit · 2026-07-23
Frontier VLMs still fail at plain-text diagrams on a new benchmark
A Reddit user reports results from ASCIITermDraw-Bench, a benchmark for getting models to draw simple diagrams in plain text. The short answer: even the strongest models still struggle with basic layout and routing.
Top results on 80 tasks
- qwen3.7-max — 77.2% ±4.0
- gemma-4-31b-it — 73.8% ±4.1
- glm-5.2 — 70.6% ±5.0
- qwen3.7-plus — 70.2% ±4.6
- deepseek-v4-pro — 68.0% ±4.7
- kimi-k2.6 — 61.8% ±6.0
- kimi-k2.7-code — 61.6% ±7.0
- minimax-m3 — 59.5% ±6.3
- nemotron-3-ultra-550b-a55b — 57.4% ±6.8
The author’s takeaway is that the best model still fails on nearly 1 in 4 tasks. Structural accuracy is relatively high, but once spacing, alignment, and routing matter, performance drops.
They argue this is not a niche academic toy problem: it’s the basic skill of expressing architecture and system design in plain text without screenshots or Mermaid, and it raises a broader question about how much “vision” in VLMs actually holds up when precision matters.
The full benchmark, methodology, and examples are linked in the post.
Related event: New Benchmark Shows Top VLMs Struggle with ASCII Diagrams(3 posts)→
More from Research
- Pi exposes cache behavior as a debate over agent harnesses burning KV caches heats up — mitsuhiko · 2026-07-23
- TPAMI paper unifies Bregman divergences with surrogate losses for learning — FrnkNlsn · 2026-07-23
- Nathan Benaich says AI will design nearly every molecule soon — nathanbenaich · 2026-07-23
- DocOps benchmark finds frontier agents still fail on long-horizon document tasks — Jiazhen Jiang · 2026-07-23
- Stanford’s vine-like soft robot grows from the tip to reach trapped people — lukas_m_ziegler · 2026-07-23
- First CAR-T Cell Therapy Approved for Solid Tumors in Gastric Cancer — Dr_Singularity · 2026-07-23