Train Qwen3-1.5B on RTX 5090 within a $5k budget
kalyan_kpl · x · 2026-09-01
A paper presents an open pretraining recipe to train Qwen3-1.5B on RTX 5090 within a $5k budget. Cost efficiency is achieved through hardware selection, low-precision training, hyperball optimization, curriculum model averaging, and data recipes. The authors released full training details for Puro-2B under Apache 2.0.
More from Infra
- Writing 100 facts into Qwen's n-gram Engram table with zero weight changes: 84% recall — Electronic_Put4530 · 2026-09-21
- Which speculative decoding setup for local Qwen models? Two llama.cpp configs compared — cradlemann · 2026-09-21
- Electron merges GPU shader disk cache fix, cutting app startup time by ~400ms — DanielLockyer · 2026-09-21
- A 3-month curated paper list for learning distributed LLM training and inference — East-Muffin-6472 · 2026-09-21
- vLLM PR adds structured generation mode for DiffusionGemma diffusion LLM — victormustar · 2026-09-21
- Anthropic's compute commitments may hit $517B over a decade, with 14.8GW already signed — FinanceYF5 · 2026-09-21