DatologyAI open-sources Zephon, cutting data-order noise from 0.82 to 0.05 points when GPU count changes

lmoroney · x · 2026-10-08

DatologyAI open-sourced Zephon, a data loader for text and multimodal training (on PyPI, with TorchTitan and Megatron-LM integrations) that keeps identical global batch order when GPU count, worker parallelism, or backend changes — even with on-the-fly tokenization, packing, and mixing. It splits data into a fixed number of logical lanes. In tests, a 1B model trained for 20B tokens showed up to 0.64 points spread on FineWeb and 0.82 on DCLM Core v1 with a standard loader depending only on GPU count; Zephon cut this to 0.011 and 0.05.

Related event: Datology AI Open-Sources Zephon, a Deterministic Streaming Dataloader(9 posts)→

Original post →

More from Infra

Infra channel →