FULL STORY
Puro-2B: Pretraining an LLM on a Single RTX 5090
Tsinghua's PACMAN team released Puro-2B, a fully open-source recipe for pretraining a 2B model from scratch on a single RTX 5090 for a few thousand dollars, drawing wide community attention.
2026-08-28 ~ 2026-08-31 · 2 episodes · 6 posts
Episode 1 · Researchers Pretrain 2B Model on a Single RTX 5090 for Under $7,000 (2026-08-28, 2 posts)
A new paper introduces Puro-2B, an open-source method that pretrains a 2B-parameter model from scratch on a single consumer RTX 5090 using FP8 precision for under $6,900, achieving performance close to Qwen2.5-1.5B.
- Puro-2B Matches Qwen2.5 Performance with $6.9K Pretraining on RTX 5090 — _reachsumit · 2026-08-28
- Puro-2B: Pretraining a 2B model from scratch on RTX 5090 GPUs for under $6.9K — iScienceLuvr · 2026-08-28
Episode 2 · Tsinghua's Puro-2B Replicates Qwen-Level Model for Under $5090 on a Single RTX 5090 (2026-08-31, 4 posts)
Tsinghua's PACMAN lab released Puro-2B, a fully open-source recipe that trains a Qwen2-1.5B-level 2B model from scratch on a single RTX 5090 for under $5090, using Blockwise FP8, the MuonH optimizer, and curriculum model averaging over 1.4T tokens.
- Train a 2B model from scratch on RTX 5090 for just $4,400 — burny_tech · 2026-08-31
- Puro-2B: Replicate Qwen2-1.5B on RTX 5090 for Under $5090 — burny_tech · 2026-08-31
- Train a Qwen2-Level Model on RTX 5090 for Under $5090 — zhaoran_wang · 2026-08-31
- Puro-2B: pretraining a 2B model on RTX 5090 for under $5090 — thu-pacman · 2026-08-31