Numerical-memory method validated on DeepSeek-V2-Lite: 2.58 MiB, 69/72 correct
Nearby_Indication474 · reddit · 2026-10-11
The author applies AKBASCORE MAM — source-free numerical memory where the model reads source text once, consolidates it into numerical memory, and answers without re-feeding the text — to a third Transformer family, DeepSeek-V2-Lite-Chat, comparing against earlier Mistral and Qwen implementations.
- Principle: append-only consolidation, frozen weights, no source text at question time
- Results across three models: Mistral-7B (93 memory tokens, 11.62 MiB, 72/72); Qwen2.5-7B (82, 4.48 MiB, 72/72); DeepSeek-V2-Lite (87, 2.58 MiB, 69/72)
- Key architectural difference: DeepSeek's MLA stores a compressed 512-dim cKV plus 64-dim rotary key instead of standard K/V caches, so its heatmap differs; dark-blue bands inside the shared 28-token instruction prefix are just low norms, not broken memory
- The 3 failures are documented in TEST597 with no verified fix; the 2.58 MiB footprint reflects MLA compression, not a measured MAM gain
- Conclusion: the mechanism works across three architectures under tested conditions, not proof of universal compatibility
- Full code, execution logs, and DOI are public
More from Research
- ETH Zurich and Google's DiskChunGS Maps Kilometer-Scale Scenes via Chunked Disk Streaming — rsasaki0109 · 2026-10-11
- GPT-6 Astra for Robotics: a curated collection of papers, reports and evaluations — NielsRogge · 2026-10-11
- 112 bugs across 84 projects: LLMs pass proof-of-concept but fail developers' own tests — lulzxdxdxd · 2026-10-11
- Lab automation only handles cookie-cutter assays — the long tail is what stifles innovation — anshulkundaje · 2026-10-11
- ArXiv caps submissions at two per month as AI paper flood overwhelms server — The Decoder · 2026-10-11
- Warning: engram/n-gram methods may markedly amplify overfitting under data repeat — SonglinYang4 · 2026-10-11