Puro-2B Matches Qwen2.5 Performance with $6.9K Pretraining on RTX 5090

_reachsumit · x · 2026-08-28

Puro-2B introduces an open, low-cost pretraining recipe that trains a 2B model from scratch on consumer-grade RTX 5090 GPUs at a compute cost of under $6.9K, approaching Qwen2.5-1.5B performance. The efficiency is achieved through hardware selection, FP8 training, hyperball optimization, curriculum model averaging, and a specific data recipe. The authors also derive a Puro Cost Scaling Law, suggesting comparable performance can be reached for around $4.4K.

Original post →

More from Infra

Infra channel →