Practice: Self-Improving AI Chip Design via RL Hits 87k tps for Kimi K3
brianryhuang · x · 2026-08-24
This blog details using Reinforcement Learning (RL) and LLMs to redesign the Kimi K3 chip, achieving full RTL-to-GDS flow automation.
Key Results:
- Phase 1: Rebuilt K3's 16x16 MAC array at 4mm², running at 1047 MHz (vs original 100 MHz).
- Phase 2: Built the full KDA dataflow accelerator, achieving 900 MHz post-route.
- Performance: The best design theoretically enables Kimi K3 to run at 87,000 tps, 10x faster than the team's example.
Technical Details:
- Explains the ASIC vs. GPU distinction, noting KDA's fixed state size makes it ideal for ASIC.
- The system is self-correcting, fixing bad SRAM synthesis iterations, forming a "Meta Harness" self-evolution loop.
- Design files are open-sourced.
More from Infra
- AI Data Centers Reshape Electricity Markets and Infrastructure — VraserX · 2026-08-24
- Crypto Miners Shift to AI for Higher Per-kWh Revenue — alvelda · 2026-08-24
- Xiaomi launches local AI host with three custom chips supporting dual models — op7418 · 2026-08-24
- Musk predicts AI and robots will explode bandwidth demand — XFreeze · 2026-08-24
- Help: Docker Deployment of Local Qwen3.8 FP8 Failing — EbbNorth7735 · 2026-08-24
- RTX 5000 Pro Runs Qwen3.8 at Unbeatable Value — Valuable-Run2129 · 2026-08-24