Same pixel-art wizard coding test: only one model finished with zero intervention
hullabaloo22 · x · 2026-09-23
FatalExit replicated majidmanzarpour's "animated pixel-art wizard in pure code" test across 4 other models. Only sol completed without any intervention; spark dropped once due to the muse code harness, Gemini failed video export twice, and Mimo got stuck and needed Codex to export. The original opus 5.5 run's prompt is shared publicly.
More from coding & agent
- theo builds his own visualizer for today's agent models, showing how cheap Luna really is — ivan_bezdomny · 2026-09-23
- Vite+ Hits RC: One Rust-Powered CLI to Replace Your Entire Web Toolchain — cnakazawa · 2026-09-23
- Tesla's in-car Grok agent books trips across Gmail, Calendar and Notion in one command — xiaohu · 2026-09-23
- Tesla's In-Car Grok Assistant Now Executes Cross-App Tasks in One Sentence — xiaohu · 2026-09-23
- Garry Tan says Capy lets him ship PRs much faster than Codex or Claude Code — garrytan · 2026-09-23
- DeskPilot: open-source native Python desktop client for local LLMs with MCP and sandboxed tools — poofph · 2026-09-23