Kimi K3 case studies show kernel optimizations, a Triton-like compiler, and a chip prototype
teortaxesTex · x · 2026-07-28
The image shows Kimi K3 case studies covering three technical areas:
- GPU kernel optimization: On both NVIDIA Hopper and an alternative-vendor GPGPU, Kimi K3 improved AttnRes latency from 283.6 ms to 114.4 ms, cut DSA and KDA runtime by 55.1% and 73.6%, and reached over half of peak TFLOPS on MLA.
- GPU compiler development: Moonshot says Kimi K3 built MiniTriton, a Triton-like compiler with custom tile-level Python frontend, MLIR-based optimization, PTX codegen, and a PyTorch-like interface. On an NVIDIA L20, it outperformed PyTorch eager and torch.compile on its benchmark suite.
- Chip design: In a 48-hour autonomous run with Kimi Code, Kimi K3 reportedly built and verified an inference-chip prototype using open-source EDA tools and the Nangate45 standard-cell library. The design targets a 4 mm² budget, near 100 MHz, and more than 8,700 tokens/s simulated decode throughput.
The page links the compiler and chip work to Moonshot’s Kimi K3 research stack, with GitHub code for minitriton and nano-kpu.
More from Infra
- Shanghai Aishengna is said to be manufacturing DUV lithography tools — zephyr_z9 · 2026-07-29
- A vendor-agnostic Vulkan backend cuts edge inference latency from 30 ms to 3 ms — ppchaos · 2026-07-29
- pdf-mcp turns technical PDFs into structured text, images, and searchable context — tom_doerr · 2026-07-29
- Moonshot’s Kimi K3 is a 2.8T open-weight MoE model with 1M-token context — alex_verem · 2026-07-29
- New scaling law paper says repetition can beat paraphrasing for some pretraining regimes — burny_tech · 2026-07-29
- SK hynix swings wildly as HBM profitability keeps rising — basedjensen · 2026-07-29