ByteDance Seed, Princeton and Tsinghua propose Falcon family for training recurrent LLM memory
burkov · x · 2026-09-06
Researchers from ByteDance Seed, Princeton, Tsinghua, UCLA and Hyperbolic tackle the long-context tradeoff: full Transformer attention is expensive as sequences grow, while recurrent models keep fixed-size memory whose update rule is itself online learning. The paper argues such memory should be trained with the representation actually available at prediction time, not same-step pairing. This yields the Falcon family: normalized memory updates with explicit control over learning speed, forgetting, and per-update window size, plus chunk-parallel GPU implementations. Experiments show competitive language modeling and better extrapolation on longer arithmetic sequences.
More from Research
- AdaptVPR: route-aware hard positive generation boosts robust visual place recognition — Shunpeng Chen · 2026-09-07
- SC Asia 2027 opens call for papers on supercomputing and AI infra, due Oct 7, 2026 — thoefler · 2026-09-07
- Frontier models double as RL teachers for smaller siblings, argues poster — haider1 · 2026-09-07
- Ex-Google Brain researcher: key algorithmic wins were found under 1e20 FLOPs, then scaled to 1e25 — _arohan_ · 2026-09-07
- GlossoGen paper: LLM agents evolve emergent languages humans can't understand — abenitezburraco · 2026-09-07
- New paper resolves three open problems in online fair division with impossibility results — chaumian · 2026-09-07