Pokédex Benchmark visually compares AI coding models with one shared prompt

kevinkern · x · 2026-09-23

Developer Kevin Kern launched the "Pokédex Benchmark": every AI coding model gets the same prompt and reference image to build a Pokédex app, with code quality scored out of 100 (reviewed by GPT-6 Astra) plus time, cost, and token stats. All apps are live and side-by-side comparable.

Key results: Opus 5.5 scored highest on code quality (80/100), while GPT-6 Astra was fastest at 8m08s but scored only 65. DeepSeek alone hit 69 in 23m27s, and a mixed Astra→DeepSeek pipeline also reached 80. The site publishes the prompt, source, speed/cost tables, and per-model notes.

Related event: Pokédex Benchmark Pits Top AI Coding Models Against One Prompt(2 posts)→

Original post →

More from coding & agent

coding & agent channel →