DatologyAI open-sources Zephon, cutting data-order noise from 0.82 to 0.05 points when GPU count changes
lmoroney · x · 2026-10-08
DatologyAI open-sourced Zephon, a data loader for text and multimodal training (on PyPI, with TorchTitan and Megatron-LM integrations) that keeps identical global batch order when GPU count, worker parallelism, or backend changes — even with on-the-fly tokenization, packing, and mixing. It splits data into a fixed number of logical lanes. In tests, a 1B model trained for 20B tokens showed up to 0.64 points spread on FineWeb and 0.82 on DCLM Core v1 with a standard loader depending only on GPU count; Zephon cut this to 0.011 and 0.05.
Related event: Datology AI Open-Sources Zephon, a Deterministic Streaming Dataloader(9 posts)→
More from Infra
- China's electricity glut turns data centers into a solution, as 14nm chips get pressed into service — teortaxesTex · 2026-10-08
- NAVER's DLoop Loops Speculative Decoding Before Verification, Gaining 5-41% Faster Inference Losslessly — naver-ai · 2026-10-08
- Transformer lead times balloon from 500 to 1,120 days, YC partner calls it a startup opportunity — ycombinator · 2026-10-08
- MIT's Christina Delimitrou uses AI to cut data center energy waste and downtime — nordicinst · 2026-10-08
- FT kicks off three-part series on China's breakneck AI infrastructure build-out, from Ulanqab to Shaoguan — zijing_wu · 2026-10-08
- Dev hits 1k tokens/sec prefill at 262K context with hybrid DeepSeek V4.1 Flash build — HankYeomans · 2026-10-08