How a 20B MoE was pre-trained for under $10/hour
const_reborn · x · 2026-07-20
A comment highlights a striking training setup from @jondurbin: a 20B MoE model was pre-trained for under $10/hour of compute.
It was not run on a cluster. Instead, the setup used eight rented single-L40S VMs spread across two continents, plus a few 4090s and a 5090, while tolerating about 6 seconds per step. The point is that pre-training — long assumed to require a large centralized cluster — can be done surprisingly cheaply with a distributed setup.
More from Infra
- Bernstein sees datacenter pipeline reaching 338 GW as AI chip demand swells — TiernanRayTech · 2026-07-21
- Super Proxy open-sources a self-hosted multi-provider LLM gateway with fallback and cost caps — Delicious-Flan88 · 2026-07-21
- Marker will get more accuracy improvements, while Chandra remains the high-accuracy option — VikParuchuri · 2026-07-21
- Nebius says SlimSpec speeds speculative decoding 8–9% without shrinking the vocabulary — Arindam_1729 · 2026-07-21
- NVIDIA brings its Cosmos 3 Edge world model to Jetson for on-device robot control — liu_mingyu · 2026-07-21
- A silicon photonic reservoir chip compensates fiber distortion in real time at 28 Gbps — bravo_abad · 2026-07-21