Q Labs: LLMs are depth-bottlenecked, loss keeps improving to 128 layers

rickasaurus · x · 2026-09-24

Q Labs Research argues depth stopped scaling: frontier models still sit near 100 layers, same as GPT-3 six years ago. Their language models keep improving to 128 layers on just 1B tokens, past frontier depths with no saturation. They reproduced contrastive RL results showing new capabilities emerging up to 256 layers (and reportedly 1,024). Scaling computational depth yields compute-efficiency gains: 1.6x at 10^20 FLOPs, projected 3.1x at 10^26. They claim current architectures cap expressive power and depth is the next scaling frontier.

Related event: Q Labs: Computational Depth Is the Missing Scaling Axis for LLMs(4 posts)→

Original post →

More from Models

Models channel →