Building CS224n's Transformer from scratch: 144 lines expose scaling and mask bugs
stanfordnlp · x · 2026-10-02
- A developer hand-built Self-Attention and Transformers following Stanford CS224n Lecture 8 (John Hewitt), passing 35/37 tests in 144 lines on the first try.
- Scaling: without the √d scale, one word put 0.999 attention weight on two words and swapping them moved the output by only 0.0003; with scaling, 0.096.
- Causal mask: a missing mask doesn't crash but drives loss to 0.037 — the model reads answers from its input. Lesson: change only critic-visible values and assert the action doesn't move.
- A 2-block decoder (127,872 params) trained on Shakespeare on a laptop for 3,000 steps: held-out loss 4.650 → 1.850.
More from Research
- HuST Lab's Multimodal Flow: Fully Continuous Unified Language-Vision Generation — hustvl · 2026-10-02
- Kuaishou's DARA Cuts Multi-Reward RL Training Steps by Up to 65% — kuaishou · 2026-10-02
- Argo-Bench Pits Data Agents Against a 7.5B-Row Warehouse; Best Model Clears Only 34.8% of Tasks — textql · 2026-10-02
- LoopCD: training-free contrastive decoding lifts looped transformers, AIME 61.9%→73.3% — arankomatsuzaki · 2026-10-02
- Is Jev secretly learning a Value function? RL calibration and System 1/2 — lateinteraction · 2026-10-02
- MLPerf Adds DLRMv4: HSTU Sequence Modeling Meets 560GB Embeddings for Production Recommenders — TheKanter · 2026-10-02