Diffusers Tensor Parallel Loading Gets Major Speedup
Hugging Face Diffusers upgraded tensor parallel loading, cutting Flux.2-Dev load times by 2.4x with TP=4 on A10G GPUs. The core PR enables shard-wise loading during .frompretrained(), saving about 89% CPU memory.
2026-10-01 ~ 2026-10-01 · 2 related posts
- Diffusers tensor parallel loading gets 2.4x faster, cuts CPU memory by 89% — RisingSayak · 2026-10-01
- Diffusers PR #14544: shard tensor-parallel checkpoints on load and save — RisingSayak · 2026-10-01