Atomic agent repo tests Kimi K3 locally against Claude Fable 5 on 3D games
testingcatalog · x · 2026-07-29
The post adds cost and token stats to the earlier atomic-agent demo and links the GitHub repo.
- The test compared Kimi K3 and Claude Fable 5 on the same prompt set for building three 3D arcade classics.
- Kimi K3 ran locally on 8× B300, produced about 237K tokens, and incurred $0 in API spend.
- The author says Kimi’s Snake scene was the best of the six outputs, with soft shadows, dust, and self-checking game logic; its maze also regenerated on restart with a live minimap and ghost tracking.
- The linked project, atomic-agent, is positioned as a local-first AI agent with long context, proper tool calling, private on-device execution, memory/skills docs, and eval folders.
Related event: Local Kimi K3 Beats Cloud Models in 3D Physics Generation Test(5 posts)→
More from coding & agent
- Non-coder builds layered memory architecture: 20k tokens tracks a year of agent conversations — matteoianni · 2026-09-11
- Warp's six non-engineering teams all run on Linear and Claude Code — mon__lim · 2026-09-11
- 9-year backend dev: AI code isn't the problem, the rate of making a mess is — Sweaty-Landscape-561 · 2026-09-11
- RTK claims token savings, but our cost benchmarks disagree — michalwarda · 2026-09-11
- Is inference latency becoming the biggest bottleneck for production AI agents? — Euphoric_Sea632 · 2026-09-11
- Anthropic researcher: 99% of engineers now run swarms of 300+ self-improving agents — AlishaOutridge · 2026-09-11