Puro-2B: pretraining a 2B model on RTX 5090 for under $5090
thu-pacman · hf · 2026-08-31
Tsinghua PACMAN's Puro-2B ("Poor Lab's Qwen2-1.5B") is a cost-efficient open-source pretraining recipe:
- Pretrains a 2B-parameter model on consumer RTX 5090 GPUs for under $5090;
- Performance nears larger baselines;
- Derives cost scaling laws and studies data curricula along the way.
Highly relevant for small teams reproducing pretraining on a budget.
Related event: Tsinghua's Puro-2B Replicates Qwen-Level Models for Under $5,000(3 posts)→
More from Infra
- Medusa from training to inference: a two-part guide to multi-token prediction acceleration — No_Progress_5399 · 2026-08-31
- OpenAI reportedly buying tens of thousands of Macs for RL training — SumitGup · 2026-08-31
- Data center reporting should learn from telegraph cable narratives — jwt0625 · 2026-08-31
- New project gpu_api introduces minimal API design, optimizing graphics rendering experience — Michael_Moroz_ · 2026-08-31
- Benchmarking Qwen3.8 BF16 across RTX 5090 and Mac via RPC — stargate425 · 2026-08-31
- Hugging Face used open weights to defend; frontier APIs failed — AlexTensor · 2026-08-31