Researchers argue computational depth is the missing scaling axis, stuck at ~100 layers since GPT-3

burny_tech · x · 2026-09-24

ethantsliu and Akshay Vegesna argue that computational depth is the missing scaling axis: params, data, sparsity, and test-time reasoning have all been scaled by orders of magnitude, but depth has been stuck at 100 layers since GPT-3.

Key claims:

This is framed as scaling computational depth per the bitter lesson; full paper details are not yet provided.

Related event: Researchers argue computational depth is the missing scaling axis(3 posts)→

Original post →

More from Research

Research channel →