Diffusers tensor parallel loading gets 2.4x faster, cuts CPU memory by 89%

RisingSayak · x · 2026-10-01

Hugging Face shipped a major upgrade to tensor parallel loading in Diffusers. On Flux.2-Dev DiT with TP degree 4 (A10G GPUs), model loading drops from 30.4s to 12.5s (2.4x faster) and per-rank peak CPU memory falls from 64.1 GB to 6.8 GB (89% less).

Previously, tensor parallelism required loading the full checkpoint on every rank before resharding. Now frompretrained() shards weights while reading from disk, delivering only each rank's slice directly into a DTensor on device, and savepretrained() is TP-aware as well. Incompatible with devicemap, quantization, flashpack, and non-safetensors formats.

Related event: Diffusers Tensor Parallel Loading Gets Major Speedup(2 posts)→

Original post →

More from Infra

Infra channel →