Turbo-dLLM Open-Sources CSBP, Speeding Diffusion LLM Training Up to 7.59x at 1M Context
Azaliamirh · x · 2026-09-19
The ScalingIntelligence team released Turbo-dLLM, an open-source distributed training library for diffusion LMs built around Context-Sharded Block Parallelism (CSBP).
- On 8x H100, DFlash2 speculative-decoder training gets 2.48x speedup at 512K and 7.59x at 1M context; block diffusion fine-tuning is 1.61x faster and AR-to-block-diffusion adaptation 1.33x faster
- Speedups grow with context length, targeting the long-context training bottleneck for agents
- Models trained with CSBP score higher on SWE-bench Verified and Terminal-Bench Lite at the same GPU hours
- Includes typed config, prepared-data runtimes, checkpointing, and optimized kernels
More from Infra
- Cloudflare Saves Another 100TB of RAM by Shrinking a Consistent Hash Ring — DanielLockyer · 2026-09-19
- YC Paper Club hosts Alternative Compute Paradigms event: optical CPUs, neuromorphics, bio GPUs — ycombinator · 2026-09-19
- HEIF Heist: one C image parser bug chain leads to RCE in OpenAI, Meta, GitHub — ccerrato147 · 2026-09-19
- AI infra market self-correcting: fake data-center projects filtered out, says weekly thesis — xtbot · 2026-09-19
- Cloudflare Saves Another 100TB of RAM With Math and Rust — f311a · 2026-09-19
- Luminal runs large-scale FLUX.2 diffusion on AMD MI300X to cut cost per image — ycombinator · 2026-09-19