Benchmarking LLMs by asking them to draw Mario on a 32x32 grid

TheMoonMidas · x · 2026-08-23

The author built pixel32bench to test various LLMs by prompting them to draw Mario on a 32x32 grid without any image examples. The experiment compares the resulting pixel art, reasoning depth, time taken, and cost across different models, serving as a window into each model's visual representation and taste.

Related event: Pixel32Bench Tests LLMs Drawing Mario on a 32x32 Grid(2 posts)→

Original post →

More from Models

Models channel →