Kimi K3 Speculative Decoding Model Hits 600K Downloads, Boosts AMD MI355X Throughput

bookwormengr · x · 2026-08-04

The community-developed Kimi-K3-DSpark speculative decoding model has surpassed 600,000 downloads on Hugging Face.

Built for Kimi K3 with support for up to 1 million tokens, the model delivers highly efficient inference. Tests on AMD MI355X hardware show a 2.2× improvement in single-stream speed and an 18% increase in peak aggregate throughput.

Technically, it extends the DFlash parallel-draft backbone with a Markov logit-bias head. Benchmarks like SWE-Rebench and GSM8K show average acceptance lengths (acclen) between 4.1 and 5.4. It maintains an acclen of 4.2 even in 1M-token contexts, drastically optimizing long-context inference efficiency.

Original post →

More from Infra

Infra channel →