Is recurrent depth the real driver? Researchers debate 3x depth vs 81x params math

xuanalogue · x · 2026-09-04

Ryan Greenblatt argues that recent gains in opaque reasoning ability are likely downstream of architecture changes, not parameter scaling. Parsing Jakub's claim of being "within a factor of 2 of GPT-4", he estimates a 3x increase in depth — which, based on open-weight scaling trends, would normally require 81x more parameters.

@xuanalogue questions the math: by default a 3x depth increase should be roughly equivalent to 3x parameters, and looped transformers actually yield less due to parameter reuse and lack of weight specialization — unless added depth enables other scaling increases.

Related event: Greenblatt: 3x Depth Gain Comes From Architecture, Not Parameter Scaling(4 posts)→

Original post →

More from Research

Research channel →