Train a 2B model from scratch on RTX 5090 for just $4,400
burny_tech · x · 2026-08-31
This paper presents an open-source pretraining recipe designed for cost efficiency on consumer-grade RTX 5090 GPUs. By combining hardware selection, Blockwise FP8, MuonH optimizer, and curriculum model averaging, the team trained a collection of Puro-2B models on up to 1.4 trillion tokens.
Key Findings:
- Cost Efficiency: The best model was trained for under $6,900. The derived "Puro Cost Scaling Law" suggests that only about $4,400 is needed to reach Qwen2-1.5B level performance.
- Performance: The model approaches the performance of Qwen2.5-1.5B.
- Methodology: The study demonstrates that full-stack co-design makes LLM pretraining accessible to small labs.
More from Infra
- ROCm 10 on dual R9 7900 boosts Qwen 27B performance by 10% — hurdurdur7 · 2026-08-31
- QQL: SQL-like language launched for Qdrant vector search simplification — qdrant_engine · 2026-08-31
- Wasmer launches one-click deploy with GitHub integration — jedisct1 · 2026-08-31
- Real-World API Cost Analysis: Token Price is a Bad Proxy, Retries and Billing Models Matter More — Few-Market-6535 · 2026-08-31
- Don't Use LLMs to Quickly Build Multi-Tenant Databases: High Maintenance Cost — kylegawley · 2026-08-31
- ClusterMAX team finds serious security holes in billion-dollar neoclouds — AccBalanced · 2026-08-31