vLLM Supports DFlash Inference Acceleration Checkpoints

vllm_project · x · 2026-07-15

Red Hat AI has released **DFlash speculative checkpoints** for two of NVIDIA's open models: - **Nemotron Ultra 550B** - **Nemotron Super 120B** Reported performance metrics include: - Average acceptance of about **5/7** draft tokens on math/reasoning tasks - Average acceptance of about **3.4/7** draft tokens on coding tasks (HumanEval) These checkpoints were trained using the open-source **Speculators** library, are licensed under **Apache 2.0**, and have been validated on the **NVIDIA B200**. Users can enable them simply by adding the `--speculative-config` parameter in vLLM.

Related event: Red Hat AI Releases DFlash Inference Checkpoints for NVIDIA Models(2 posts)→

Original post →

More from Infra

Infra channel →