Fine-tuning framework comparison surfaces EP, SP and 2M-sequence training details

StasBekman · x · 2026-07-26

A detailed comparison of popular fine-tuning frameworks is described as “slightly misguided but still very useful.”

The post highlights practical caveats around parallelism settings: it says EP cannot be combined with CP, but ALST/SP works fine with EP; because SP overlaps with DP, you can run EP=8 plus SP=8 on a single node and still train a 2M-sequence model, even a 30B MoE, on 8× H200s. The main complaint is that many published “max” numbers fail to state whether they assume one GPU or many, or what GPU class was used.

Related event: Fine-Tuning Framework Comparison Sparks Parallelism Debate(3 posts)→

Original post →

More from Infra

Infra channel →