Researchers argue computational depth is the missing scaling axis, claiming LLMs are depth-bottlenecked

burny_tech · x · 2026-09-23

Researchers argue computational depth is the missing scaling axis: params, data, sparsity, and test-time reasoning have all been scaled by orders of magnitude, but depth has stayed at 100 layers since GPT-3.

They claim LLMs are severely depth-bottlenecked and bad at using the depth they have, and that architectural interventions lifting this bottleneck yield gains that grow with compute.

Related event: Researchers Argue Computational Depth Is a Missing Scaling Axis for LLMs(2 posts)→

Original post →

More from Research

Research channel →