Storing model-native numerical memory outside a frozen LLM: replicated from Qwen to Mistral, 127/128 Top-1
Nearby_Indication474 · reddit · 2026-10-07
After completing AKBASCORE MAM on Qwen2.5-7B-Instruct, the author independently localized and replicated the same memory mechanism on Mistral-7B-Instruct-v0.3: 127/128 correct memories at Top-1 (99.22%), counterfactual retrieval 125/128 (97.66%), shifted-pointer control 0/128. No weight changes, no fine-tuning, no LoRA, no learned router, and no gold memory ID is supplied at retrieval time.
Key ideas:
- Model-native numerical memory: source text passes through the frozen transformer once; the K/V states from its own attention are extracted as a numerical memory structure, rather than storing human-readable text for the model to reread.
- External persistence: the numerical structure can live in RAM or be serialized to disk while weights stay frozen.
- Endogenous attention pointer: with the active numerical A-memory installed in the transformer cache, the frozen model's own attention channel (L28H00 on Mistral, dubbed ÇAĞRIİZ) produces a question-conditioned distribution over A-memory positions — not an external router or trained classifier, just native Q·K computation.
- K vs V roles: K answers "where to look"; V constructs a 1024-dim (8 KV heads × 128 dims, L00-V uncentered) question-conditioned retrieval address, then all 128 B-memories compete via max cosine similarity — no filenames or memory IDs involved.
The author argues the deeper question than 99.22% is what exactly the model is reading: this path stores machine-native numerical states derived from the model itself, diverging from conventional text/embedding retrieval.
More from Models
- Temp 0 doesn't guarantee deterministic LLM output — batch shape and kernels shift logits — JFPuget · 2026-10-07
- OpenAI doesn't need more resets, needs better communication, argues paying user — AirportEither2456 · 2026-10-07
- GPT-6-luna Unlocks More Reasoning Tokens via API: ~18k Tokens Scores ~80.5% on Terminal-Bench — LysandreJik · 2026-10-07
- Bug Hunt Benchmark retest: GPT-6.1 Sol recovers, Muse still cheapest strong agent — PawelHuryn · 2026-10-07
- Subscription multipliers recalculated: SuperGrok gives 80x, Muse Power+Contributor 137x usage — PawelHuryn · 2026-10-07
- r/mathematics Reacts to OpenAI's New Math Proofs, Takes Are Split — petburiraja · 2026-10-07