REMORY: Learning Residual Memory for Context Compaction
Hanchen Xia, Baoyou Chen, Yutang Ge, Naihao Deng, Senqiao Yang, Zilong Dong, Weihao Yuan, Siyu Zhu
cs.CL, cs.AI
2026-10-08
Soft memory tokens appended after the compaction summary lift frozen-LLM agent scores up to 9.8 points and nearly match full-context SummHay at 5.2% of input positions.
Long-horizon agents live on a sawtooth: input grows with interaction, collapses to a summary at compaction, then grows again. /compact in Claude Code and Codex CLI works exactly this way, and the summary becomes the checkpoint the agent resumes from.
The loss is real and unpredictable. Keeping the main thread does not guarantee the summary supports every later decision, and what counts as relevant from earlier context shifts as the task unfolds. The paper's position is blunt: you cannot enumerate in advance what the summary will be missing, so writing better summaries is not a fix.
Human cognition is the reference point. Working memory holds a few chunks, yet people sustain projects for months by drawing on long-term memory; hippocampal indexing theory has partial cues activate an index that reinstates earlier experience. REMORY gives the agent a learned neural index alongside the explicit summary.
At each compaction boundary, given a freshly generated summary, a memory network maps the current history into a sequence of soft tokens, at most 4,096 vectors living directly in the frozen LLM's input embedding space, concatenated after the summary. The summary is the readable checkpoint; the tokens carry whatever else subsequent prediction needs. Because the tokens are conditioned on the summary and appended after it, the arrangement works like a residual connection along the sequence dimension: summary as the trunk, memory as the residual.
Architecturally, every 1,024 source positions produce 64 soft tokens, hierarchically compressed to the 4,096 cap while preserving source order. The network uses frozen actor features plus learned queries for parallel prediction, with gated cross-attention to the summary. The Qwen variant builds on DFlash (1.94B trainable parameters); the GLM variant uses DFlash-2 (1.502B). The actor stays frozen throughout; gradients pass through it to train only the memory network.
Training runs in two stages, teacher and student sharing the same frozen actor:
The loss is reverse KL from student to teacher over the full vocabulary and response. The Stage II design matters: since the summary already exists, memory only has to close the remaining prediction gap, not reconstruct every omitted fact. Training diagnostics in Figure 3 show reconstruction initialization converging better than random init under the same objective.
SummHay, 92 queries, Qwen3.8-27B, with actor, query, decoding, and base summary held fixed:
| Representation | Input positions | Coverage | Citation F1 | Joint |
| Raw context | 96,872 | 72.46 | 56.60 | 43.41 |
| Summary | 1,393 | 67.55 | 56.98 | 39.84 |
| Summary++ (equal-budget text) | same budget | 67.52 | 56.94 | 39.78 |
| Summary + REMORY | 5,024 | 67.95 | 61.03 | 43.39 |
Two comparisons carry the section. First, summary++ is indistinguishable from summary alone: spending the same extra positions on generated text buys nothing, while soft memory adds 4.04 citation F1 and 3.55 joint score (paired bootstrap 95% interval +0.74 to +6.21). Second, at 5.2% of raw-context positions, the memory run matches the raw-context joint score (43.39 vs 43.41), with coverage 4.5 points lower. Attribution gains far exceed coverage gains: memory mostly helps connect recovered claims back to evidence.
Agent benchmarks are all within-actor comparisons with weights, Codex CLI harness, prompts, tools, decoding, budgets, and compaction policy fixed; both sides can retrieve exact history:
| Benchmark | Qwen3.8-27B w/o → w/ | GLM-5.3-Flash w/o → w/ |
| AutomationBench | 35.5 → 45.3 | n/a |
| JobBench | 33.4 → 41.0 | n/a |
| Terminal-Bench 2.1 | 71.9 → 76.4 | 84.3 → 87.6 |
| BrowseComp | 74.0 → 77.0 | 84.9 → 89.0 |
AutomationBench and JobBench average five runs (the +9.8 on AutomationBench is the largest single gain); Terminal-Bench and BrowseComp are single runs. Across the four actor-benchmark pairs, repeated tool outputs drop 17.9–29.8% and tool errors 25.3–49.7%, consistent with better use of earlier feedback. Generation cost falls everywhere except JobBench ($3.09 to $3.21). In 488 blind pairwise judgments of the first post-compaction action, the memory-backed action wins 70.7%.
The Live Trading case is the most eye-catching and the weakest: two parallel Qwen paper-trading accounts, ROI −11.86% without memory versus +5.11% with it, generation spend $303.81 versus $170.30. The authors themselves note the trajectories differ in market selection and timing, so the return gap is directional only.
Every team building long-horizon agents has hit the wall where compaction loses information, and the only industrial answer today is writing better summaries. This paper establishes two things: at equal position budget, a learned continuous representation preserves attribution cues far better than generated text, and the gains carry through to end-to-end agent scores and tool-use quality. The three-way split is pragmatic: readable checkpoint, learned supplement, and retrieval for exact records all coexist, so nothing about the current harness has to be torn up.
The barrier is equally clear: each actor needs its own 1.5B–1.9B memory network, and a model upgrade means retraining. Teams that can run training get a usable recipe; API-only users get a paper for now. The claim that residual memory pushes the 320B-parameter GLM-5.3-Flash toward frontier level rests on the uncontrolled Figure 1 comparison; the controlled evidence is the within-actor deltas.
Stated by the authors: Live Trading covers only two divergent trajectories; the cell-segmentation case cannot prove the restored details came from memory rather than re-reading files; cost estimates cover token-priced generation only, excluding encoding and memory-network execution; the frontier-model scores in Figure 1 use their original harnesses and are not controlled comparisons.
Beyond that: Terminal-Bench and BrowseComp are single runs, so 3–4 point gains carry no variance information. The memory is bound to one actor's input space, trained separately for Qwen and GLM, with no cross-model transfer tested. Continuous memory is not human-readable, so it cannot be audited the way a summary can. SummHay coverage still trails raw context by 4.5 points: compaction stays lossy, and memory restores attribution more than coverage.