Maglev: Sliding Recurrent Memory Solves Transformer Forgetting Issue

anselm · x · 2026-08-19

Maglev introduces a recurrent Transformer architecture that balances full history attention with computational efficiency using a sliding-window mechanism.

Core Mechanism:

Results: Outperforms sliding-window and latent recurrent Transformer baselines on validation loss and pretraining benchmarks.

Original post →

More from Research

Research channel →