Researchers argue LLMs are severely depth-bottlenecked: the missing scaling axis
madhavsinghal_ · x · 2026-09-23
The industriaalist team argues computational depth is the missing scaling axis: params, data, sparsity and test-time reasoning have all scaled by orders of magnitude, yet LLM depth has stayed at 100 layers since GPT-3. Their findings: LLMs are severely depth-bottlenecked and bad at using the depth they have, and architectural interventions that lift this bottleneck yield gains that grow with compute. Reposter madhavsinghal notes principled depth scaling will likely recur in architecture autoresearch that works.
More from Research
- Continuous diffusion beats discrete on random k-SAT, proposed as standard benchmark — ArashVahdat · 2026-09-23
- Grady Booch: Contemporary AI Still Lacks Abductive Reasoning, Just 'Next-Token Prediction' — Grady_Booch · 2026-09-23
- AI solves Navier-Stokes-related problem as machines upend mathematics, New Scientist reports — burny_tech · 2026-09-23
- Mathematician says OpenAI likely proved a significant partial case of the Hodge conjecture — burny_tech · 2026-09-23
- Code benchmarks are mostly slop: dev calls for narrow evals per domain, not one score — almmaasoglu · 2026-09-23
- LaBinius: lattice-based polynomial commitments hit tens of thousands of hash proofs per second — burny_tech · 2026-09-23