Jon Durbin says he pre-trained a 20B MoE for under $10 an hour
const_reborn · x · 2026-07-29
A repost of a DropZone episode where Jon Durbin explains how he pre-trained a 20B MoE for under $10 per hour of compute.
Key ideas mentioned in the episode:
- a surrogate model 48× smaller than the real expert stands in for it
- workers train on data they never directly see
- ternary weights change the hardware requirements
- existing sync algorithms failed, so he wrote a new one
The post points readers to the full episode and notes the project is tied to $TAO.
More from Infra
- How to serve spiky trillion-token LLM workloads with autoscaling and multi-region routing — zainhas · 2026-07-29
- AI research is splitting between one RTX 3090 and NVL72-scale clusters — _xjdr · 2026-07-29
- Ridges launches x402 on Ridgeline, letting AI agents pay for coding infrastructure — bittingthembits · 2026-07-29
- Stripe opens agent stablecoin payments in 27 more countries — jeff_weinstein · 2026-07-29
- Flow Traders picks CoreWeave as its primary AI cloud platform — _ScottCondron · 2026-07-29
- H100 rental prices firm to $2.75/hour as compute market tightens — BenBajarin · 2026-07-29