Long-read dissects activation functions: how slope, curvature and saturation shape gradients and trainability
CatAstro_Piyush · x · 2026-10-09
Inspired by a tweet joking that "activation functions are eating your jobs," this long-form essay mathematically dissects these one-line, zero-parameter rules that nonetheless decide whether a network trains at all.
- Removing an activation collapses a hundred layers into one linear map; changing its shape can make the same architecture easy or slow to train, sparse, smooth, or unstable.
- The author builds a "skeleton" anatomy: shoulders = input domain, skull = output range, spine = forward transform ϕ(z), elbows = slope (derivative controlling gradient flow), knees = curvature (how slope changes), plus saturation and parameters.
- The "brain" part explains how these pieces affect forward signals, backward gradients, representation geometry, and trainability — contrasting ReLU chopping the negative half, Sigmoid compressing the real line into a probability interval, GELU/SiLU leaking small negative signals, and the never-settling sine.
- Core claim: differences that look cosmetic on a plot all enter training.
More from Research
- Neural Radiance Caching variant speeds up specular lighting in real-time path tracing — ssh4net · 2026-10-09
- ABC releases open stack for scalable behavior cloning with real and sim teleoperation data — rsasaki0109 · 2026-10-09
- How Those Morphogenesis Simulations Work: Deformable Meshes With Controllable Fibers — zzznah · 2026-10-09
- OpenAI's frontier model produces new math results, including an 84-page proof of the Erdős–Pomerance conjecture — burny_tech · 2026-10-09
- Researcher: activation monitors beat black-box monitoring for AI cyber safety — burny_tech · 2026-10-09
- ETH's SpaceFlow: training-free locally controllable 3D generation from text and primitives — ethz · 2026-10-09