Why sequence padding bloats training systems: pipeline parallelism is to blame
青稞AI · wechat · 2026-09-12
A deep-dive post-mortem on why post-training frameworks become slow and inelegant, tracing the rot to pipeline parallelism and sequence padding. Covers the bubble-fighting lineage (GPipe to DualPipe), why RL post-training's dynamic inputs break fixed schedules (RLHFuse, PipelineRL, StreamRL), why padding to [B,S] is unnecessary given FlashAttention varlen and Orca-style packing, torchtune's FlexAttention workaround and its complexity/memory costs, and real bugs like Unsloth's gradient-accumulation loss error and Megatron-Core's FLOPs overestimation caused by padding.
More from Infra
- Intel's silicon photonics couplers hit 1-1.5 dB IL, with visible epoxy delamination flaws — jwt0625 · 2026-09-12
- CPO paper criticized for vague DLW-to-PIC coupling description: 'such as TCB' — jwt0625 · 2026-09-12
- During AWS outage, one engineer kept enterprise services up with just 22 min downtime — generativist · 2026-09-12
- US hosts 43% of global datacenter power use; China just 13%, report finds — TMWNN · 2026-09-12
- Intel engineers show wafer-level chiplet testing for co-packaged optics paper — jwt0625 · 2026-09-12
- Intel engineers spotted testing chiplets and packages on wafer-level tester — jwt0625 · 2026-09-12