New Gated Recurrent Transformer Cuts Params and Memory by 60%

burkov · x · 2026-08-30

A paper from Cerebras, GUC, and TUM introduces a gated recurrent transformer architecture. It uses state-conditioned modulation across iterated shared layers to match standard language model performance while reducing parameters and peak decoding memory by roughly sixty percent.

Related event: Gated Recurrent Transformer Matches GPT-2 with Far Fewer Parameters(2 posts)→

Original post →

More from Research

Research channel →