vLLM Supports DFlash Inference Acceleration Checkpoints
vllm_project · x · 2026-07-15
Red Hat AI has released **DFlash speculative checkpoints** for two of NVIDIA's open models: - **Nemotron Ultra 550B** - **Nemotron Super 120B** Reported performance metrics include: - Average acceptance of about **5/7** draft tokens on math/reasoning tasks - Average acceptance of about **3.4/7** draft tokens on coding tasks (HumanEval) These checkpoints were trained using the open-source **Speculators** library, are licensed under **Apache 2.0**, and have been validated on the **NVIDIA B200**. Users can enable them simply by adding the `--speculative-config` parameter in vLLM.
Related event: Red Hat AI Releases DFlash Inference Checkpoints for NVIDIA Models(2 posts)→
More from Infra
- UK AI datacentres face backlash over heat, noise and land use — nordicinst · 2026-07-21
- Fluidstack raises $830M at $7.5B valuation as Anthropic backs a $50B compute buildout — rohanpaul_ai · 2026-07-21
- Early Krea2 Gradio WebUI targets 6GB low-VRAM local runs — Fluid_Kaleidoscope17 · 2026-07-21
- Z.AI starts running a 1GW AI data center built entirely on domestic chips — Polymarket · 2026-07-21
- Local models feel far more capable once paired with the right harness — Soft-Barracuda8655 · 2026-07-21
- Voice-agent teams should use platforms first, then own STT events when failures get weird — FollowingSuitable941 · 2026-07-21