Gated Recurrent Transformer Matches GPT-2 with Far Fewer Parameters
Cerebras, together with researchers from GUC and TU Munich, proposed the Gated Recurrent Transformer (GRT), which reuses a small shared Transformer core with hidden-state gating to decouple depth from parameter count, matching GPT-2 with roughly half the parameters.
2026-08-30 ~ 2026-08-30 · 2 related posts
- GRT Architecture: Halves Params to Match GPT-2 Performance — burny_tech · 2026-08-30
- New Gated Recurrent Transformer Cuts Params and Memory by 60% — burkov · 2026-08-30