Diffusers Tensor Parallel Loading Gets Major Speedup

Hugging Face Diffusers upgraded tensor parallel loading, cutting Flux.2-Dev load times by 2.4x with TP=4 on A10G GPUs. The core PR enables shard-wise loading during .frompretrained(), saving about 89% CPU memory.

2026-10-01 ~ 2026-10-01 · 2 related posts