Maglev: Sliding Recurrent Memory improves long-context efficiency
UTEXAS · hf · 2026-08-15
UTEXAS introduced Maglev, a recurrent Transformer architecture with a fixed-size memory. By utilizing coupled prefiller-decoder training, the method improves long-context modeling performance while enabling efficient parallel training and reducing inference costs.
More from Research
- TalkRL: Optimizing insulin infusions with Clinical RL — tw_killian · 2026-08-15
- Proposal: Buy Research Data Directly for Pretraining in Exchange for Compute — teortaxesTex · 2026-08-15
- SuperFlex Introduces Bending and Tapering for Improved Superquadric Point Cloud Decomposition (ECCV 2026) — CSProfKGD · 2026-08-15
- FLIM imaging reveals long-distance non-neural bioelectric patterns — drmichaellevin · 2026-08-15
- Open source closes the gap with closed labs: Quality gap now just months — togethercompute · 2026-08-15
- AI-assisted research trends spark a surge in NeurIPS submissions and quality advances — PTenigma · 2026-08-15