KVMem pages agent context overflow to KV state, beating compaction on DeepSWE

omarsar0 · x · 2026-09-09

Elvis Saravia highlights a new paper, KVMem, targeting inference efficiency for long-running agents whose workspaces outgrow the context window long before the task finishes.

Existing approaches fall short: compaction loses fine-grained execution evidence, while text retrieval re-prefills content the model already processed.

KVMem's approach:

Results: on the DeepSWE long-context test with Qwen3.8-27B, task success rises from 43.8% under compaction to 48.4%.

Local deployment also stands out — it runs Qwen3.6/3.8-27B NVFP4 with MTP on a laptop with a 24GB RTX 5090 while virtualizing a much longer context.

Original post →

More from coding & agent

coding & agent channel →