Cerebras paper: layer dropout saves up to 25% training FLOPs and yields 1.55x faster decoding

burny_tech · x · 2026-09-09

A Cerebras research paper, "Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference," argues that layer dropout should return to state-of-the-art pretraining recipes, countering the belief that it hurts accuracy at scale.

Method

Findings

All pretraining experiments ran on Cerebras CS-3 systems.

Original post →

More from Infra

Infra channel →