xLLM open-sourced: flexible pre-training infra hits 10,050 tokens/s/GPU on H200 without dataset rebuilds

HongyiWang10 · x · 2026-09-29

IFM AI released xLLM, an open-source infrastructure for pre-training and fine-tuning dense and MoE LLMs that keeps key training decisions changeable — tokenizer, data mixture, architecture, and stages — without rebuilding the dataset. On H200s it achieves 6,295 tokens/s per GPU on K2-Horizon-MoVA-36B-A4B and 10,050 on Llama3-8B. Ships with K2 Horizon checkpoints, logs, and recipes.

Original post →

More from Infra

Infra channel →