Diffusers tensor parallel loading gets 2.4x faster, cuts CPU memory by 89%
RisingSayak · x · 2026-10-01
Hugging Face shipped a major upgrade to tensor parallel loading in Diffusers. On Flux.2-Dev DiT with TP degree 4 (A10G GPUs), model loading drops from 30.4s to 12.5s (2.4x faster) and per-rank peak CPU memory falls from 64.1 GB to 6.8 GB (89% less).
Previously, tensor parallelism required loading the full checkpoint on every rank before resharding. Now frompretrained() shards weights while reading from disk, delivering only each rank's slice directly into a DTensor on device, and savepretrained() is TP-aware as well. Incompatible with devicemap, quantization, flashpack, and non-safetensors formats.
Related event: Diffusers Tensor Parallel Loading Gets Major Speedup(2 posts)→
More from Infra
- turbopuffer rewrites storage engine v3 in public, grinding query plans live — vboykis · 2026-10-01
- BofA Raises Micron Targets on Durable AI Memory Cycle, Sees $60-100B Buyback by FY28-29 — firstadopter · 2026-10-01
- Distilled 260M text-to-image model hits <190ms on a T4, generating per keystroke — Gradio · 2026-10-01
- Agent spins up 322 Hugging Face Jobs in 90 minutes to test code, total bill ~$4 — vanstriendaniel · 2026-10-01
- Anthropic reportedly buying ~5GW of compute from Broadcom, may become its largest customer by 2027 — zephyr_z9 · 2026-10-01
- Supertonic: open-source 99M-param local TTS beats ElevenLabs in tests — JafarNajafov · 2026-10-01