Why sequence padding bloats training systems: pipeline parallelism is to blame

青稞AI · wechat · 2026-09-12

A deep-dive post-mortem on why post-training frameworks become slow and inelegant, tracing the rot to pipeline parallelism and sequence padding. Covers the bubble-fighting lineage (GPipe to DualPipe), why RL post-training's dynamic inputs break fixed schedules (RLHFuse, PipelineRL, StreamRL), why padding to [B,S] is unnecessary given FlashAttention varlen and Orca-style packing, torchtune's FlexAttention workaround and its complexity/memory costs, and real bugs like Unsloth's gradient-accumulation loss error and Megatron-Core's FLOPs overestimation caused by padding.

Original post →

More from Infra

Infra channel →