SlideDP fine-tunes Qwen2.5-72B on four RTX 4090s, beating FSDP2 throughput by 11.2%
Ruijia Yang · hf · 2026-10-01
- Problem: Host-resident layer streaming enables full-parameter LLM fine-tuning beyond GPU memory, but data-parallel ranks compete for shared host resources; replicated transfers amplify traffic and strong scaling exposes host work.
- Approach: SlideDP is a synchronous data-parallel runtime for shared-host multi-GPU systems: one authoritative host state, communication routes decoupled from state layout, and pipelining of parameter delivery, gradient aggregation, and CPU updates across ranks and chunks, guided by an analytical step-time model.
- Results: Geometric-mean throughput ratios of 1.46–2.64x over SlideFormer, MegaTrain, and ZeRO-Offload. On four H100s it approaches GPU-resident FSDP2 throughput for Qwen3-14B; with larger batches it processes over 1M tokens per step, exceeding FSDP2's measured peak by 11.2%. It also supports 256K-token sequences and fine-tunes Qwen2.5-72B on four RTX 4090 GPUs. Code open-sourced.
More from Infra
- Huawei Mate 90 debuts Kirin 9050 Pro, first chip with LogicFolding architecture — ingliguori · 2026-10-01
- NVIDIA A20 Standard Reportedly Skips WMCM Packaging — 'Not Enough Capacity' — zephyr_z9 · 2026-10-01
- TRL hits 1M post-trainings per month, team eyes 1M per week — QGallouedec · 2026-10-01
- oMLX 0.7.0 lands with up to 50% faster long-context decoding and a rebuilt memory guard — pcuenq · 2026-10-01
- China sets all-time monthly electricity record, but industrial power demand growth slows — teortaxesTex · 2026-10-01
- NVIDIA Dynamo Snapshot cuts LLM serving cold-start times by ~10x — Nasereliver · 2026-10-01