Stanford's CS336 course draws learners deep into SwiGLU, from ReLU's dead neurons to modern LLMs
stanfordnlp · x · 2026-09-26
A learner shares their experience with Stanford's CS336 "Language Modeling From Scratch" YouTube playlist as a way to build strong AI/ML foundations.
- The course prompted a deep dive into activation functions: ReLU, SwiGLU (Swish Gated Linear Unit), and how SwiGLU fixes ReLU's "dead neuron" problem
- Fun fact: the original 2017 Transformer base/large architectures used ReLU in their FFN layers, while modern models like Llama, Llama 2, and Mistral use SwiGLU
More from Companies & People
- Jab at Meta: its last homegrown hit was Marketplace, quips X user — yungcontent · 2026-09-26
- OpenAI hires Snowflake's Americas sales SVP, doubling down on enterprise GTM — thedealdirector · 2026-09-26
- Zuckerberg Says Brute-Forcing Compute Is the Path to ASI, Signals Bigger Meta CapEx — RihardJarc · 2026-09-26
- Altman: OpenAI is conducting an extensive review of its agents' internet use during training, HF incident most severe — sama · 2026-09-26
- GCC Bans AI-Assisted Code While LLVM Allows It — But You Must Still Understand Your Code — lemire · 2026-09-26
- Fireship recaps Meta Connect 2026: Meta is pivoting again — Fireship · 2026-09-26