New Gated Recurrent Transformer Cuts Params and Memory by 60%
burkov · x · 2026-08-30
A paper from Cerebras, GUC, and TUM introduces a gated recurrent transformer architecture. It uses state-conditioned modulation across iterated shared layers to match standard language model performance while reducing parameters and peak decoding memory by roughly sixty percent.
Related event: Gated Recurrent Transformer Matches GPT-2 with Far Fewer Parameters(2 posts)→
More from Research
- IBM Proposes Spatial Matryoshka Training for Multi-Granularity Document Retrieval — _reachsumit · 2026-09-01
- Doc-REFRAG: coarse-compress then selectively expand for faster, more accurate multi-image RAG — _reachsumit · 2026-09-01
- Google Releases RSLM: Training-Free Vector Quantization for ANN Search — _reachsumit · 2026-09-01
- CHAP Framework Enables Personalized Generative Retrieval with Single-Pass Inference — _reachsumit · 2026-09-01
- Alibaba Jointly Trains Embeddings and Codebooks for E-commerce Generative Retrieval — _reachsumit · 2026-09-01
- Alibaba Proposes PAO to Prevent Embedding Collapse in RL Fine-Tuning for Retrieval — _reachsumit · 2026-09-01