2,400 experiments show layer dropout can match dense baselines in LLM pretraining

burkov · x · 2026-09-10

A systematic study revisits layer dropout—randomly skipping transformer blocks during training—long abandoned due to reported accuracy loss at scale. Researchers ran 2,400+ controlled experiments on decoder-only transformers from 271M to 8.2B parameters, training on up to 160B tokens with Cerebras CS-3 systems, jointly optimizing the configuration to see if layer dropout can match or beat dense baselines while cutting training FLOPs and inference latency.

Related event: Cerebras paper: layer dropout cuts training compute by 25%(2 posts)→

Original post →

More from Infra

Infra channel →