How DatologyAI Generated 12 Trillion Synthetic Tokens — And Fixed 4 Pipeline Bottlenecks
AI Engineer · youtube · 2026-10-03
DatologyAI CTO Bogdan Gaza shares engineering lessons from generating 12 trillion synthetic tokens (web, math, code) for frontier pre-training, since the web offers only 30T usable tokens.
- BeyondWeb: seeded rephrasing beats generating synthetic data from scratch.
- Stack moved from split Slurm/Kubernetes to unified Ray, KubeRay and vLLM on EKS on HyperPod.
- Four bottlenecks and fixes: batched S3 metadata fetches (11 days → 2 hours), right-sized partitions plus checkpointing for GPU failures, cross-cluster CPU/GPU scheduling, and vLLM flag sweeps yielding 40% more inference throughput.
More from Infra
- Two renderers share one L40S: voxlap ported to CUDA Rust streams multiplayer views in real time — idanbeck · 2026-10-03
- Anthropic to spend at least $518B on AI infrastructure over a decade, IPO filing shows — Beth_Kindig · 2026-10-03
- LibLayaX: run the Laya decision model in your app at 670 decisions/sec, no server needed — felipedaragon · 2026-10-03
- A local/remote LLM router dies as new model releases outpace its training — gaviniboom · 2026-10-03
- INT21: 2 engineers direct AI to build 20 inference engines in 2 weeks, up to 2.4× faster than SGLang — bingxu_ · 2026-10-03
- Traversal's 5 Levels of Self-Driving Production: Why Coding Agents Make Ops Harder — AI Engineer · 2026-10-03