Red Hat AI Releases vLLM Checkpoints to Accelerate Inference
vllm_project · x · 2026-07-16
Red Hat AI has released DFlash speculator checkpoints for two open-source NVIDIA LLMs:
- Nemotron Ultra 550B
- Nemotron Super 120B
Trained on vLLM's open-source Speculators library under the Apache 2.0 license, they are validated on NVIDIA B200. The official vLLM activation parameter is: --spec-model ... --spec-tokens 7 --spec-method dflash.
Performance-wise, math and reasoning tasks accept an average of 5 out of roughly 7 draft tokens, while coding tasks (HumanEval) accept an average of 3.4 out of 7.
Related event: Red Hat AI Releases DFlash Inference Checkpoints for NVIDIA Models(2 posts)→
More from Infra
- Vercel AI Gateway data shows Anthropic, OpenAI and Google at 97.09% spend share — cramforce · 2026-07-21
- NVIDIA starts rolling out 102.4 Tbps Spectrum-6 switches for Vera Rubin AI factories — nvidia · 2026-07-21
- Mustafa Suleyman says Microsoft is preparing for an OpenAI exit, while a new chip costs 30% less than GB200 — thoefler · 2026-07-21
- Microsoft and Mistral sign multi-billion-dollar deal to expand AI infrastructure in Europe — The Decoder · 2026-07-21
- Speculative decoding boosts Qwen3.6-27B on one 5090, but slows crowded servers — luke_pacman · 2026-07-21
- NVIDIA says Blackwell Ultra hit 1,648 TFLOPs per GPU on DeepSeek-V3 671B training — NVIDIAAI · 2026-07-21