TokenSpeed stabilizes Kimi K3 on ~1,000 B300 GPUs for 3+ weeks: GB300 TP8 optimization details

zhyncs42 · x · 2026-09-08

Open-source LLM inference engine TokenSpeed published Part I of its Kimi K3 on GB300 optimization series (TP8 within a single NVL72 NVLink domain), now stable for 3+ weeks in production across 1,000 B300 and 2,000 H200 GPUs, with engineering support and GB300 NVL72 access from NVIDIA.

Key results: using multi-turn SWE-Smith traces as the workload and EAGLE3 speculative decoding (lightseekorg/kimi-k3-eagle3-mla), they studied concurrency 1–16; EAGLE3 improves the throughput-vs-speed curve across the range; communication-aware sharding of the two LatentMoE projections reclaims 7.7 GiB of parameters per GPU while cutting tail latency; replay-based KDA state recovery shrinks verify-time workspace from 16.7 GiB to 0.4 GiB per GPU at maxnumseqs=64. Parts II (expert parallelism) and III are coming.

Original post →

More from Infra

Infra channel →