Kimi K3 Wins in Arcade Game Generation Test
OfirPress · x · 2026-07-17
A hands-on comparison reveals that Kimi K3 performed the best on a task to create three retro arcade games, while also costing less.
Comparison data:
- Kimi K3: 18.4K tokens, $0.28
- GPT-5.6: 18.1K tokens, $0.28
- Opus 4.8: 21.3K tokens, $0.54
The tester noted that the Qbert game built by Kimi K3 was the best of the three. Its Battle City was also the closest to the original, with correct logic for tank movement, aiming, shooting, and wall destruction. While GPT-5.6 had the best visual style, it failed in Battle City, with broken tank and base logic.
Related event: Kimi K3 Tops Frontend Code Arena and Sparks Debate(53 posts)→
More from Models
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11