xLLM open-sourced: flexible pre-training infra hits 10,050 tokens/s/GPU on H200 without dataset rebuilds
HongyiWang10 · x · 2026-09-29
IFM AI released xLLM, an open-source infrastructure for pre-training and fine-tuning dense and MoE LLMs that keeps key training decisions changeable — tokenizer, data mixture, architecture, and stages — without rebuilding the dataset. On H200s it achieves 6,295 tokens/s per GPU on K2-Horizon-MoVA-36B-A4B and 10,050 on Llama3-8B. Ships with K2 Horizon checkpoints, logs, and recipes.
More from Infra
- Dagger founder: the Great CI Bottleneck of 2026 is a software problem, not hardware — msharmas · 2026-09-29
- LayerSkip: Self-Speculative Decoding Speeds Up LLMs Without a Draft Model — burkov · 2026-09-29
- SpaceX outlines supercomputer, Terafab, Gigasat and new Louisiana Starbase plans — elonmusk · 2026-09-29
- Musk says space will hold nearly all compute; Google tests if TPUs work there — CackleRooster · 2026-09-29
- 95+ TPS and 262K context for Qwen 27B on a single RTX 3090 with LlamAmpere v0.4 — Brief-Tap-6616 · 2026-09-29
- More Americans oppose a local data center than a nuclear reactor, says Cathie Wood — PeterDiamandis · 2026-09-29