Drawing the Mona Lisa: GPT-5.6 Beats Claude, Gemini, and Grok

ArtificialOther · x · 2026-08-12

TryAI built a "drawing arena" to test the capabilities of four frontier vision models—GPT-5.6, Claude Fable 5, Grok 4.5, and Gemini 3.6 Flash—using only colored pencil tools on a blank canvas.

The testing included objective image reproduction (like the Mona Lisa) and open-ended text prompts. The system meticulously tracked every stroke, tool-usage cost, and the final output.

The authors note that fuzzy, open-ended tasks provide a more intuitive visual indicator of capability than standard benchmarks, effectively highlighting the gap between frontier and open-weight models. In this experiment, GPT-5.6 delivered the best performance, while Grok 4.5 struggled significantly with the basic drawing task.

Original post →

More from Fun

Fun channel →