FULL STORY

Puro-2B: Pretraining an LLM on a Single RTX 5090

Tsinghua's PACMAN team released Puro-2B, a fully open-source recipe for pretraining a 2B model from scratch on a single RTX 5090 for a few thousand dollars, drawing wide community attention.

2026-08-28 ~ 2026-08-31 · 2 episodes · 6 posts

Episode 1 · Researchers Pretrain 2B Model on a Single RTX 5090 for Under $7,000 (2026-08-28, 2 posts)

A new paper introduces Puro-2B, an open-source method that pretrains a 2B-parameter model from scratch on a single consumer RTX 5090 using FP8 precision for under $6,900, achieving performance close to Qwen2.5-1.5B.

Episode 2 · Tsinghua's Puro-2B Replicates Qwen-Level Model for Under $5090 on a Single RTX 5090 (2026-08-31, 4 posts)

Tsinghua's PACMAN lab released Puro-2B, a fully open-source recipe that trains a Qwen2-1.5B-level 2B model from scratch on a single RTX 5090 for under $5090, using Blockwise FP8, the MuonH optimizer, and curriculum model averaging over 1.4T tokens.