MemFold: Fixed-Budget Soft Memory Tops PersonaMem via On-Policy Optimization
Jingxuan Wu · hf · 2026-10-03
MemFold introduces a fixed-budget soft memory for long-horizon personalized assistants, optimizing memory by the behavior it supports rather than text reconstruction:
- Method: a query-conditioned textual memory is compressed into K continuous vectors forming the reader's memory interface; the reader is trained on its own rollouts with group-relative rewards for task outcomes plus confidence-gated on-policy distillation from a frozen textual-memory teacher that never samples and is removed at inference.
- Motivation: retaining information differs from acting on it; text memory grows input length while latent memories are typically trained to reconstruct text or imitate reference answers on sequences the reader never produced.
- Results: across three Qwen backbones, MemFold attains the highest accuracy on PersonaMem-32K and PersonaMem-128K with widening margins at longer history, and transfers to PrefEval and LongMemEval without target-domain training.
- Ablations: most gain comes from the reward term, with a smaller boost from the teacher signal; memory interventions show the reader relies on instance-specific soft-memory content.
More from Research
- CMU blog: harness engineering boosts scientific agents without retraining models — niloofar_mire · 2026-10-03
- Automated agent harness search vs human taste: no winner across drug design tasks — niloofar_mire · 2026-10-03
- AI agents settle all 15,973 semigroups of order 6 with 5M lines of verified Lean — KyleCranmer · 2026-10-03
- Dev open-sources Peacebell, a from-scratch 291M WWII domain LLM built over 11 months — wayneworkman · 2026-10-03
- Dev vibecodes AtlasBench Europe spatial reasoning benchmark; GPT-6.1 tops at 84.67% — flowersslop · 2026-10-03
- Light-powered AI detects deepfakes with nearly 98% accuracy — ai-edition · 2026-10-03