Layer dropout works: decay over training, drop top layers more, enables speculation

yaroslavvb · x · 2026-09-11

yaroslavvb shares hands-on findings that layer dropout works when you (1) start with high dropout and decay it, (2) drop top layers more often than bottom layers, and (3) tune hyperparameters carefully. Properties 1-2 approximate stochastic growing, and a bonus is that the kept (likely bottom) layers can act as an independent draft model for speculative decoding. He also describes a mobile paper-reading workflow: a 10-minute audio summary plus GPU-running agents answering clarifying questions while he walks.

Related event: Engineer shares mobile paper-reading workflow and layer dropout findings(2 posts)→

Original post →

More from Research

Research channel →