Karpathy Highlights Simon Willison's 'Pelicans on Bicycles' LLM Eval Test
karpathy · x · 2026-08-03
Andrej Karpathy shared Simon Willison's keynote from the AI Engineer World's Fair, uploading the source code so it's playable and forkable in the browser.
Willison noted that with over 30 significant models released recently, traditional benchmarks and leaderboards are losing trust. He increasingly relies on his own eval: asking text-only LLMs to generate an SVG of a pelican riding a bicycle. This method tests the models' ability to understand and generate complex code, proving to be a surprisingly practical evaluation tool.
More from Models
- DeepSeek V4 Flash Performance Varies Wildly Across Coding Agents — PMinervini · 2026-08-03
- Comparing 33 Qwen Models: Over 1,100 One-Shot Outputs Analyzed — kms_dev · 2026-08-03
- Rumor: Zhipu's GLM 5.5 and DeepSeek Pro Set for August Release — bindureddy · 2026-08-03
- MiniMax H3 Multimodal Video Model Coming to ComfyUI — NerdyRodent · 2026-08-03
- OpenAI's Prepaid API Credits Expire After One Year, Sparking Developer Outrage — mark_k · 2026-08-03
- The Benchmaxxing Plague: Expert Breaks Down AI Eval Flaws — AI Engineer · 2026-08-03