ReLU Paper Review: Solving the Vanishing Gradient Problem

burkov · x · 2026-08-15

Stanford professor Andrej Karpathy shared a review of the papers on ReLU and Dropout. He argues that ReLU sparked a greater revolution in neural networks than Dropout because the vanishing gradient problem made deep networks with more than two hidden layers practically untrainable before ReLU, whereas overfitting could be mitigated by techniques like early stopping.

Original post →

More from Research

Research channel →