Diffusers PR #14544: shard tensor-parallel checkpoints on load and save
RisingSayak · x · 2026-10-01
The follow-up to Diffusers' tensor parallelism support: PR #14544 by JingyaHuang makes frompretrained(..., parallelconfig=...) shard weights while reading, so each rank loads only its own slice straight into a DTensor on device, and savepretrained() is TP-aware. Unsupported combos (devicemap, quantization, flashpack, DDUF, non-safetensors) raise errors.
Related event: Diffusers Tensor Parallel Loading Gets Major Speedup(2 posts)→
More from Infra
- Classic CUDA tutorial walks from naive matmul to 94% of cuBLAS performance step by step — Abhishekcur · 2026-10-01
- Micron sees memory supply tighter through 2028, with 26 take-or-pay deals and $150B RPO — firstadopter · 2026-10-01
- Oracle, the largest 'Chinese Cloud' in America, surges on 100% RPO growth forecast — kevinsxu · 2026-10-01
- Raspberry Pi raises 2GB Pi 4/5 prices by $12.50 as memory costs keep climbing — jedisct1 · 2026-10-01
- Turso Cloud expands to Australia, Brazil, Canada and Sweden, now in 10 AWS regions — glcst · 2026-10-01
- Full AI Chat and Image Generation on a 1990 Tandy 286: 9s Draw Times, Open Source — jacobpederson · 2026-10-01