Jon Barron walks through backpropagation by hand on a tiny two-layer network
techNmak · x · 2026-10-06
Jon Barron demonstrates backpropagation on a minimal network (one input, two ReLU hidden units, linear output). A forward pass yields a loss of 0.5; the chain rule then propagates gradients backward through the computation graph, combining upstream gradients with local derivatives and summing contributions across paths. Both pre-activations are positive, so ReLU slopes are 1 and gradients pass through unattenuated—negative pre-activations would zero the gradient for that path. After computing all four weight gradients at the same original parameter values, a plain SGD step drops the loss to 0.12005 (with the caveat that arbitrary learning rates don't guarantee loss reduction). Key takeaway: backprop computes the gradients, SGD changes the weights—and the math scales unchanged to real networks.
More from Research
- UT Austin math chair: OpenAI appears set to release ~400 AI-generated proofs at once — 141_1337 · 2026-10-06
- RT-SAFE benchmark: frontier VLMs hit 94.1% task success but only 0.7% finish safely — Lianhuiq · 2026-10-06
- Dynamic weight grafting localizes how LLMs store facts learned during finetuning — ChenhaoTan · 2026-10-06
- Strogatz to Wolfram: your Prisoner's Dilemma tournament stopped before evolution kicked in — stevenstrogatz · 2026-10-06
- 0.08 correlation on million-dollar datasets: researchers question virtual cell capital allocation — anshulkundaje · 2026-10-06
- What do LLMs actually mean when they say they're uncertain? A calibration debate — sineadwilliamso · 2026-10-06