Single 8x 5090 Rig Hits 167k tokens/s Training Throughput, Beating DDP
jon_durbin · x · 2026-07-30
The author demonstrated a training throughput of 167k tokens/s on a single machine equipped with 8x 5090 GPUs, using the Parallax framework on a toy MoE model (5B total / 1B active parameters).
Key performance and optimization metrics:
- High MFU: Achieved an equivalent 60% Model FLOPs Utilization under BF16.
- Better Loss Curve: Beats the DDP (Distributed Data Parallel) baseline at matched steps.
- Efficient Convergence: Reached 3.068 nats at 19B tokens (Llama-3 tokenizer). Compared to the DDP baseline, it reaches the token budget 15% faster and 64% cheaper.
The author speculates this performance gap is largely due to using separate Muon optimizers for experts, incorporating surrogate feedback and fold normalization. They are now preparing for larger-scale training runs.
Related event: 8x5090 Single Node Achieves 167k tokens/s Training Throughput(2 posts)→
More from Infra
- The Race for Power: Assessing Global Electricity Production for AI — lemire · 2026-07-30
- Musk Reveals xAI Infrastructure: Minihard and Macroharder Pack 220k GB300s — kevinnbass · 2026-07-30
- Down $600M in a Day: Inside Leopold's AI Infrastructure Investment Thesis — ivan_bezdomny · 2026-07-30
- YC Paper Club Dives into Multi-GPU Kernels, Inference Efficiency and Heterogeneous Hardware — Y Combinator · 2026-07-30
- SK Hynix Earnings Analysis: AI Memory Demand Strong, Market Overreacts to Oversupply — tengyanAI · 2026-07-30
- Buildcleaner reclaims 443GB of disk space by cleaning build artifacts, free and open-source MIT — jasonkneen · 2026-07-30