Pokédex Benchmark visually compares AI coding models with one shared prompt
kevinkern · x · 2026-09-23
Developer Kevin Kern launched the "Pokédex Benchmark": every AI coding model gets the same prompt and reference image to build a Pokédex app, with code quality scored out of 100 (reviewed by GPT-6 Astra) plus time, cost, and token stats. All apps are live and side-by-side comparable.
Key results: Opus 5.5 scored highest on code quality (80/100), while GPT-6 Astra was fastest at 8m08s but scored only 65. DeepSeek alone hit 69 in 23m27s, and a mixed Astra→DeepSeek pipeline also reached 80. The site publishes the prompt, source, speed/cost tables, and per-model notes.
Related event: Pokédex Benchmark Pits Top AI Coding Models Against One Prompt(2 posts)→
More from coding & agent
- Veteran dev says GPT-6 Astra is the first coding model he fully trusts — josh_bickett · 2026-09-23
- LLM coding tip: always specify what sits behind a paywall, or the app stays free — jdluk87 · 2026-09-23
- Opus 5.5 builds guitar store sim with 300+ playable guitars that turns into a beat 'em up — chongdashu · 2026-09-23
- Coco MCP 0.2.0: Native MCP debugger now renders massive responses instantly — camiloazula · 2026-09-23
- A Gemini agent to auto-reset your 50+ leaked passwords: a killer use case — sup_nim · 2026-09-23
- OpenAI startup engineering lead: in 2026 'everything is a coding agent' — simple and elegant wins — RichmanRonald · 2026-09-23