DeepLoop paper makes looped transformers scalable; rumor claims frontier models are 48 layers looped twice

StartupYou · x · 2026-09-03

yifanzhang claims 'some frontier models are basically a 48-layer transformer looped twice (48L x 2)'—unverified—and introduces DeepLoop: Depth Scaling for Looped Transformers, a paper making looped transformer training stable and scalable. If true, frontier labs may be reusing layers rather than merely stacking depth, but the rumor remains unconfirmed.

Original post →

More from Models

Models channel →