Open-source project trains an LLM from scratch in PyTorch, 13M params on free Colab
thisguyknowsai · x · 2026-09-29
A fully open-source, step-by-step project takes you from raw data to a model that writes English: download and prep The Pile (825GB of text), hand-code multi-head attention, MLP blocks, layer norm, and causal masking in PyTorch straight from "Attention Is All You Need," then train and generate. 13M params trains on a free Google Colab GPU; up to 2B params on a single A100; correct English in under a day on the free tier. Every code section comes with an explanation.
More from coding & agent
- Scraping Xiaohongshu hit posts with Codex + a wired Android phone — huangyun_122 · 2026-09-29
- Relic turns multi-agent collaboration failures into executable org protocols, +5.7pp delivery — ulab-ai · 2026-09-29
- Xiaomi's GAGAR: quality-aware advantage redistribution improves code agent RL — XiaomiMiMo · 2026-09-29
- Same cyberpunk Oregon Trail prompt, wildly different build times on two AI tools — BertMacklenF8I · 2026-09-29
- Do separate verification agents actually fix AI coding's false 'it works'? — T_hompson · 2026-09-29
- ITIS MCP Server Brings the Taxonomic ITIS Database to LLMs via Model Context Protocol — modelcontextprotocol · 2026-09-29