DFlash Accelerates Blackwell Inference
PyTorch · x · 2026-07-10
PyTorch introduced DFlash, a PyTorch-based block-diffusion model for speculative decoding aimed at accelerating LLM inference on NVIDIA Blackwell.
The post notes that researchers released a related paper and implementation details: DFlash uses a lightweight block-diffusion drafter to generate candidate tokens in parallel, which are then verified by the target model. NVIDIA states that at the same interaction level, it can boost the throughput of gpt-oss-120b by up to 15x, and it is now available in TensorRT-LLM, SGLang, and vLLM.
More from Infra
- Gritt raises a new round to automate solar array installation and maintenance — rebeccakaden · 2026-07-21
- Refactoring 150k LOC Takes 96 Hours: Is Compute the Bottleneck for AI Coding? — robleclerc · 2026-07-21
- Seeking Recommendations: Essential Local Small Models (Audio/Vision/TTS) — DeepOrangeSky · 2026-07-21
- Compute Allocation Limits: The Root Cause of Missing Architecture Innovation in European LLMs — IgorCarron · 2026-07-21
- AI Energy Footprint Pales Compared to Transport and Agriculture — dreamwieber · 2026-07-21
- SkyPilot comes out of stealth with claims of 10x faster AI time-to-intelligence — songhan_mit · 2026-07-21