Train a 135M parameter LM in 2.5 hours: A full journey

EAccelerate_42 · x · 2026-08-20

A developer documented the complete journey of training a tiny 135M parameter language model in just 2.5 hours. The workflow covers Continued Pretraining (CPT), Supervised Finetuning (SFT), Preference Optimization (DPO), and Reinforcement Learning (GRPO). The tech stack includes data generation, Unsloth, Hugging Face, evaluation harness creation, and structured output generation.

Original post →

More from Models

Models channel →