Layer dropout works: decay over training, drop top layers more, enables speculation
yaroslavvb · x · 2026-09-11
yaroslavvb shares hands-on findings that layer dropout works when you (1) start with high dropout and decay it, (2) drop top layers more often than bottom layers, and (3) tune hyperparameters carefully. Properties 1-2 approximate stochastic growing, and a bonus is that the kept (likely bottom) layers can act as an independent draft model for speculative decoding. He also describes a mobile paper-reading workflow: a 10-minute audio summary plus GPU-running agents answering clarifying questions while he walks.
Related event: Engineer shares mobile paper-reading workflow and layer dropout findings(2 posts)→
More from Research
- RD-Forget: reversible, query-dependent forgetting for agent memory — MaryamMiradi · 2026-09-11
- Science Advances editor: no author has ever disclosed AI use despite policy — TuhinChakr · 2026-09-11
- YOCO explained: one shared KV cache reused across the model's second half — stochasticchasm · 2026-09-11
- PARSER: parallel chunk subagents with an RL-trained lead agent for long-context QA — omarsar0 · 2026-09-11
- Year-long study: heavier AI companion engagement predicts lower well-being — dhadfieldmenell · 2026-09-11
- DeepMind launches AlphaGenome Atlas, a 1TB navigable map of human DNA — neil_chilson · 2026-09-11