Jon Barron walks through backpropagation by hand on a tiny two-layer network

techNmak · x · 2026-10-06

Jon Barron demonstrates backpropagation on a minimal network (one input, two ReLU hidden units, linear output). A forward pass yields a loss of 0.5; the chain rule then propagates gradients backward through the computation graph, combining upstream gradients with local derivatives and summing contributions across paths. Both pre-activations are positive, so ReLU slopes are 1 and gradients pass through unattenuated—negative pre-activations would zero the gradient for that path. After computing all four weight gradients at the same original parameter values, a plain SGD step drops the loss to 0.12005 (with the caveat that arbitrary learning rates don't guarantee loss reduction). Key takeaway: backprop computes the gradients, SGD changes the weights—and the math scales unchanged to real networks.

Original post →

More from Research

Research channel →