Training Apex-2 for ~$2,000: dual GH200 + DiLoCo and 57% duplicate data lessons
Prestigious-Taste-63 · reddit · 2026-10-05
An indie developer shared lessons from training Apex-2 for about $2,000:
- GPUs were the bottleneck: planned for 1T tokens, managed only 80B; a single H100 fell far short (FineWeb-Edu alone is 1.3T tokens).
- DiLoCo setup: instead of a 2x H100 instance ($4.19/GPU/hour), used two GH200 instances ($2.29/GPU/hour), each hitting 40% MFU, merging every 350 steps (usually under a minute) — about 1.9x faster than one GPU and cheaper.
- Dedup: MinHash dedup across the full corpus found 57% of the FineWeb-Edu sample and 34% of DCLM were duplicates, since FineWeb-Edu only dedups within single Common Crawl snapshots. The author notes the FineWeb team reported no gains from cross-snapshot dedup — a trade-off, not a free win.
- Architecture: followed the Mixtral paper for MoE.
The author is now training a 21B MoE model and hopes to join a lab.
More from Infra
- SEA compute buildout: GPUs arrive faster than data centers, land and power are the bottleneck — vincent_koc · 2026-10-05
- Rumor: SPCX reportedly developing its own optical fiber — Kyrannio · 2026-10-05
- llama.cpp merges new Metal kernels, making speculative decoding 3.4x faster on M3 Ultra — ggerganov · 2026-10-05
- Whistle, an on-device speech recognition model, trends on Hugging Face — Cactus-Compute · 2026-10-05
- Test equipment and laser lead times stretch up to a year as AI demand surges — zephyr_z9 · 2026-10-05
- Qualcomm licenses Huawei's LogicFolding chip tech in multiyear cross-license deal — teortaxesTex · 2026-10-05