Train a 2B model from scratch on RTX 5090 for just $4,400

burny_tech · x · 2026-08-31

This paper presents an open-source pretraining recipe designed for cost efficiency on consumer-grade RTX 5090 GPUs. By combining hardware selection, Blockwise FP8, MuonH optimizer, and curriculum model averaging, the team trained a collection of Puro-2B models on up to 1.4 trillion tokens.

Key Findings:

Related event: Tsinghua's Puro-2B Replicates Qwen-Level Model for Under $5090 on a Single RTX 5090(4 posts)→

Original post →

More from Infra

Infra channel →