Cerebras paper: layer dropout cuts training compute by 25%
A Cerebras paper with roughly 2,400 experiments shows that properly configured layer dropout matches dense baselines in LLM pretraining while cutting training compute by about 25% and speeding up inference by 1.55x, challenging the belief that dropout hurts large-scale pretraining.
2026-09-09 ~ 2026-09-10 · 2 related posts
- Cerebras paper: layer dropout saves up to 25% training FLOPs and yields 1.55x faster decoding — burny_tech · 2026-09-09
- 2,400 experiments show layer dropout can match dense baselines in LLM pretraining — burkov · 2026-09-10