Puro-2B Matches Qwen2.5 Performance with $6.9K Pretraining on RTX 5090
_reachsumit · x · 2026-08-28
Puro-2B introduces an open, low-cost pretraining recipe that trains a 2B model from scratch on consumer-grade RTX 5090 GPUs at a compute cost of under $6.9K, approaching Qwen2.5-1.5B performance. The efficiency is achieved through hardware selection, FP8 training, hyperball optimization, curriculum model averaging, and a specific data recipe. The authors also derive a Puro Cost Scaling Law, suggesting comparable performance can be reached for around $4.4K.
More from Infra
- Miles now supports RL training for Qwen and GLM models with high-performance kernels — ying11231 · 2026-08-28
- AI Semiconductor Endgame 2026: Infra Bubble and Open vs Closed Economics — sudoraohacker · 2026-08-28
- GLM-5.3-Flash hits 270 tok/s: 10% higher quality than 5.2 at one-tenth the cost — Yuchenj_UW · 2026-08-28
- Ornith-1.5-35B-A3B runs agentic coding at 32 tok/s on an 8GB RTX 3070 laptop — Elemental_Particle · 2026-08-28
- From MySQL to Redis: Internet Scaling History and AI Lessons — generativist · 2026-08-28
- Intel XE3P projected specs: 1.3 PFLOPS FP8, 1.5TB/s bandwidth, 2027 launch — QuixiAI · 2026-08-28