How a 20B MoE was pre-trained for under $10/hour

const_reborn · x · 2026-07-20

A comment highlights a striking training setup from @jondurbin: a 20B MoE model was pre-trained for under $10/hour of compute.

It was not run on a cluster. Instead, the setup used eight rented single-L40S VMs spread across two continents, plus a few 4090s and a 5090, while tolerating about 6 seconds per step. The point is that pre-training — long assumed to require a large centralized cluster — can be done surprisingly cheaply with a distributed setup.

Original post →

More from Infra

Infra channel →