Researchers argue LLMs are severely depth-bottlenecked: the missing scaling axis

madhavsinghal_ · x · 2026-09-23

The industriaalist team argues computational depth is the missing scaling axis: params, data, sparsity and test-time reasoning have all scaled by orders of magnitude, yet LLM depth has stayed at 100 layers since GPT-3. Their findings: LLMs are severely depth-bottlenecked and bad at using the depth they have, and architectural interventions that lift this bottleneck yield gains that grow with compute. Reposter madhavsinghal notes principled depth scaling will likely recur in architecture autoresearch that works.

Original post →

More from Research

Research channel →