Kimi K3 Wins Blind Test for Code Rendering
MaziyarPanahi · x · 2026-07-18
A same-prompt blind test was conducted on Kimi K3, GLM-5.2, and Opus 4.8: they were asked to generate a complete 3D world using only a single HTML file and three.js. A blinded Qwen3-VL then scored the rendering effects without knowing which model produced which output.
Kimi K3 took first place, ranking at the top across multiple worlds. Opus 4.8 followed closely and excelled in the "Emerald City" scenario. GLM-5.2 ranked third but maintained structural stability without breaking down in any world. The author specifically noted that while Kimi K3's open weights won't be released until July 27, it already outperforms closed-source frontier models on such creative rendering tasks.
Related event: Kimi K3 Stuns with Coding and 3D Reasoning, Beating SOTA Models(5 posts)→
More from Models
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Claude is no longer available for minors as Anthropic rolls out age assurance — Muhammad523 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11