Benchmarking LLMs by asking them to draw Mario on a 32x32 grid
TheMoonMidas · x · 2026-08-23
The author built pixel32bench to test various LLMs by prompting them to draw Mario on a 32x32 grid without any image examples. The experiment compares the resulting pixel art, reasoning depth, time taken, and cost across different models, serving as a window into each model's visual representation and taste.
Related event: Pixel32Bench Tests LLMs Drawing Mario on a 32x32 Grid(2 posts)→
More from Models
- Long-horizon coding training makes models talk to themselves, not to you — alejandroll10 · 2026-08-23
- OpenAI's Daybreak Blue model reportedly refuses to audit user's own project — AIandDesign · 2026-08-23
- User Feedback: Opus 5 Feels Like a Significant Downgrade from 4.6 — kimmonismus · 2026-08-23
- Claude Pricing Transparency Criticized: Why Silicon Valley Is Hated — StewartalsopIII · 2026-08-23
- RTX 5090 runs Qwen3.8-27B at 262K context — Fz1zz · 2026-08-23
- Flashback: GPT-4 cost $60/M output tokens with 8K context three years ago — gajesh · 2026-08-23