ByteDance Seed, Princeton and Tsinghua propose Falcon family for training recurrent LLM memory
burkov · x · 2026-09-06
Researchers from ByteDance Seed, Princeton, Tsinghua, UCLA and Hyperbolic tackle the long-context tradeoff: full Transformer attention is expensive as sequences grow, while recurrent models keep fixed-size memory whose update rule is itself online learning. The paper argues such memory should be trained with the representation actually available at prediction time, not same-step pairing. This yields the Falcon family: normalized memory updates with explicit control over learning speed, forgetting, and per-update window size, plus chunk-parallel GPU implementations. Experiments show competitive language modeling and better extrapolation on longer arithmetic sequences.
More from Research
- AI decodes whale 'talk' by probing the latent space of their calls — maier_ak · 2026-09-07
- New paper explores Conformal Prediction for offensive security attacks — chaumian · 2026-09-07
- HiSfM: scaffold-anchored hierarchical SfM tames repeated-structure ambiguity and cuts runtime — zhenjun_zhao · 2026-09-07
- BLASt3R (ECCV'26): uncalibrated bundle adjustment beats all prior calibrated VSLAM methods — zhenjun_zhao · 2026-09-07
- 4-month-old Chinese startup Atomelody debuts Melo-1, beating AlphaFold3 on most benchmarks — 新智元 · 2026-09-07
- Economists keep getting overparameterization wrong: SGD's implicit regularization is the point — Afinetheorem · 2026-09-07