A colored-pencil benchmark finds vision models can draw, but not stop at their best frame

新智元 · wechat · 2026-07-25

TryAI ran a colored-pencil drawing arena where four vision models had to paint Mona Lisa, Starry Night, and five open-ended prompts one stroke at a time using the same tool set.

What the experiment tested

Main findings

Cost and workflow differences

The broader takeaway is that these models could see their own work, but lacked a reliable “stop” decision. In a multi-step creative task, knowing when a result is already good enough became as important as planning or execution.

Original post →

More from Multimodal

Multimodal channel →