Puro-2B: Pretraining a 2B model from scratch on RTX 5090 GPUs for under $6.9K

iScienceLuvr · x · 2026-08-28

A new arXiv paper, Puro-2B, presents an open, cost-efficient pretraining recipe reproducible on consumer hardware. The authors trained from scratch on RTX 5090 GPUs with FP8 precision over 1.4T tokens — 22,514 active GPU-hours, $6.9K compute cost, 17.6 elapsed days — with the best checkpoint surpassing Qwen2-1.5B and approaching Qwen2.5-1.5B.

The paper highlights the reproduction cost barrier: training Llama-3.2-3B costs over $1.5M, and reproducing SmolLM3-3B over $700K. Efficiency comes from hardware selection, low-precision training, hyperball optimization, curriculum model averaging, and the data recipe. The authors also derive a "Puro Cost Scaling Law" relating training cost to average performance, and open-source the full pipeline covering data recipe, infrastructure, software and training strategy.

Related event: Researchers Pretrain 2B Model on a Single RTX 5090 for Under $7,000(2 posts)→

Original post →

More from Infra

Infra channel →