Researchers argue computational depth is the missing scaling axis, stuck at ~100 layers since GPT-3
burny_tech · x · 2026-09-24
ethantsliu and Akshay Vegesna argue that computational depth is the missing scaling axis: params, data, sparsity, and test-time reasoning have all been scaled by orders of magnitude, but depth has been stuck at 100 layers since GPT-3.
Key claims:
- Forcing a 100,000-bit hidden state into 17-bit text tokens for chain-of-thought is a massive bottleneck and a wasteful human prior
- LLMs are both severely depth-bottlenecked and bad at using the depth they have
- Architectural interventions that efficiently lift this bottleneck yield gains that grow with compute
This is framed as scaling computational depth per the bitter lesson; full paper details are not yet provided.
Related event: Researchers argue computational depth is the missing scaling axis(3 posts)→
More from Research
- Omnii genome language model research preview debuts on Latent Space podcast — JosephJacks_ · 2026-09-24
- Static word embeddings reproduce the 'LLM pain axis' result, undercutting claims of felt experience — wschroll · 2026-09-24
- CLM-8B: contrastive System One model claims 9x faster inference, agentic SOTA — ChengleiSi · 2026-09-24
- Tencent ARC releases GAE: geometry-native latents halve camera error, cut FVD up to 23% — CSProfKGD · 2026-09-24
- OpenRSI founders on why RSI is the missing piece of ASI, open to everyone — ChengleiSi · 2026-09-24
- OpenRSI-Index calls for domain leads and compute partners to scale open RSI — ChengleiSi · 2026-09-24