TokenSpeed stabilizes Kimi K3 on ~1,000 B300 GPUs for 3+ weeks: GB300 TP8 optimization details
zhyncs42 · x · 2026-09-08
Open-source LLM inference engine TokenSpeed published Part I of its Kimi K3 on GB300 optimization series (TP8 within a single NVL72 NVLink domain), now stable for 3+ weeks in production across 1,000 B300 and 2,000 H200 GPUs, with engineering support and GB300 NVL72 access from NVIDIA.
Key results: using multi-turn SWE-Smith traces as the workload and EAGLE3 speculative decoding (lightseekorg/kimi-k3-eagle3-mla), they studied concurrency 1–16; EAGLE3 improves the throughput-vs-speed curve across the range; communication-aware sharding of the two LatentMoE projections reclaims 7.7 GiB of parameters per GPU while cutting tail latency; replay-based KDA state recovery shrinks verify-time workspace from 16.7 GiB to 0.4 GiB per GPU at maxnumseqs=64. Parts II (expert parallelism) and III are coming.
More from Infra
- MagicaLabs says $4M pretraining beat all public base models, 50x more compute-efficient than DeepSeek — magicailabs · 2026-09-09
- Qualcomm confirms AWS deal is baked into its $15B FY29 data center revenue target — BenBajarin · 2026-09-08
- Data center construction spend jumps $25B in six months, job openings top 300k — AccBalanced · 2026-09-08
- CDOs backed by GPU leases likely coming as compute financialization accelerates — AccBalanced · 2026-09-08
- Citi TMT takeaways: hyperscaler backlog near $1.7T, AI capex tracking to $3.8T by 2030 — sanjaykalra · 2026-09-08
- AI data center interconnect chip startup Celero raises $275M at $3B+ valuation — dinabass · 2026-09-08