Cerebras paper: layer dropout cuts training compute by 25%

A Cerebras paper with roughly 2,400 experiments shows that properly configured layer dropout matches dense baselines in LLM pretraining while cutting training compute by about 25% and speeding up inference by 1.55x, challenging the belief that dropout hurts large-scale pretraining.

2026-09-09 ~ 2026-09-10 · 2 related posts