Open-source project trains an LLM from scratch in PyTorch, 13M params on free Colab

thisguyknowsai · x · 2026-09-29

A fully open-source, step-by-step project takes you from raw data to a model that writes English: download and prep The Pile (825GB of text), hand-code multi-head attention, MLP blocks, layer norm, and causal masking in PyTorch straight from "Attention Is All You Need," then train and generate. 13M params trains on a free Google Colab GPU; up to 2B params on a single A100; correct English in under a day on the free tier. Every code section comes with an explanation.

Original post →

More from coding & agent

coding & agent channel →