Drawing the Mona Lisa: GPT-5.6 Beats Claude, Gemini, and Grok
ArtificialOther · x · 2026-08-12
TryAI built a "drawing arena" to test the capabilities of four frontier vision models—GPT-5.6, Claude Fable 5, Grok 4.5, and Gemini 3.6 Flash—using only colored pencil tools on a blank canvas.
The testing included objective image reproduction (like the Mona Lisa) and open-ended text prompts. The system meticulously tracked every stroke, tool-usage cost, and the final output.
The authors note that fuzzy, open-ended tasks provide a more intuitive visual indicator of capability than standard benchmarks, effectively highlighting the gap between frontier and open-weight models. In this experiment, GPT-5.6 delivered the best performance, while Grok 4.5 struggled significantly with the basic drawing task.
More from Fun
- Doing RL with LLMs is like rediscovering human nature from first principles — burny_tech · 2026-08-12
- Counter-intuitive AI Agent Behavior: Stop Extrapolating Human Thinking onto Agents — krishnan · 2026-08-12
- Opus 5 Autonomously Builds GTA6 in 24 Hours: City Generation & Self-Management — ukanwat · 2026-08-12
- Google Search Epic Fail: Sam Altman Declared "Dead" — MrOuzo · 2026-08-12
- Feeling AI Agents Through a 'Sixth Sense': The Developer's Symbiosis Experience — DionysianAgent · 2026-08-12
- HuggingFace Joins Hype: Qwen3.8-27B Release Draws Massive Anticipation — huggingface · 2026-08-12