WarpState: 2B-param architecture swaps growing KV cache for fixed-size dual associative memory
zemondza · reddit · 2026-09-21
An experimental 2B-parameter architecture, WarpState, replaces ever-growing token-level KV states with a compact recurrent memory: two associative banks per head (fast/slow decays, β≈0.90/0.99) updated by bounded normalized EMA writes, plus exact causal attention within 128-token chunks. Streaming state stays 49 MiB regardless of stream length — though the author stresses stream length ≠ effective memory horizon.
Related event: WarpState: Two-Timescale Memory to Replace KV Cache(2 posts)→
More from Research
- Tiny-task benchmarks no longer make sense for frontier models, argues dev — pvncher · 2026-09-22
- Why collaborating agents beat subagents: broadcasting breakthroughs in parallel — IgorCarron · 2026-09-22
- Frank Noe Clarifies How Their AI Solves the Electronic Schrödinger Equation — FrankNoeBerlin · 2026-09-22
- Did OpenAI Solve the Wrong Navier-Stokes Problem? Scientific American Weighs In — scientificamerican · 2026-09-22
- Harvard open sources LLM inference traces: 6.12 billion requests from a year of production traffic — markjeffrey · 2026-09-22
- Zhejiang U & SJTU unveil DAS, an agent that writes publication-ready surveys in an hour — jiqizhixin · 2026-09-22