Single 8x 5090 Rig Hits 167k tokens/s Training Throughput, Beating DDP

jon_durbin · x · 2026-07-30

The author demonstrated a training throughput of 167k tokens/s on a single machine equipped with 8x 5090 GPUs, using the Parallax framework on a toy MoE model (5B total / 1B active parameters).

Key performance and optimization metrics:

The author speculates this performance gap is largely due to using separate Muon optimizers for experts, incorporating surrogate feedback and fold normalization. They are now preparing for larger-scale training runs.

Related event: 8x5090 Single Node Achieves 167k tokens/s Training Throughput(2 posts)→

Original post →

More from Infra

Infra channel →