IFM Lab releases xLLM training library: hot-swappable tokenizers and 1k-line Jinja chat templates

HildeKuehne · x · 2026-09-29

IFM Lab open-sourced xLLM, a training library built around a flexible online data pipeline for fine-tuning on chat/instruction data: tokenizers and data mixtures stay changeable during training, so a new tokenizer or chat template is just a config change, not a reprocessing run. Features include tokenization straight from JSONL, configurable mixtures with buffered shuffle, bestfit packing that keeps documents intact, and throughput unaffected by overlapping CPU prep with GPU training. The team says flexibility mattered in practice — they kept finding bugs in chat templates (their final Jinja is 1k lines) and only fast tokenizer/template swaps let them fix things mid-training.

Related event: IFM Lab Open-Sources xLLM Training Framework with Hot-Swappable Tokenizer and Data Mix(2 posts)→

Original post →

More from Infra

Infra channel →