Pre-registered study: selection beats extraction in agent memory, 3061x cheaper to write
Rishabh Sharma · hf · 2026-10-07
This pre-registered study asks whether conversational memory needs LLM-extracted facts, or if selecting the right raw turns suffices — prior published results conflict.
Key findings from LoCoMo and LongMemEval experiments:
- At tight budgets, raw turns selected by a single call to a typed decision model (Jev) are non-inferior to LLM-extraction memory (one-sided 95% bound -3.0 vs a -5-point margin); blind human grading narrows but doesn't change the result
- Raw turns cost 3061x less to write; the result holds with a second answer model
- Reranking gains shrink as budget grows: +17.4 on LoCoMo / +9.1 on LongMemEval keeping 3 of 30 candidates, but only +1.5 / +1.1 at generous budgets, where extraction is more accurate — explaining why published results disagree
- At matched context, Jev selects as accurately as an LLM reranker (bound -2.0) at a third of the latency, beating multi-call graph traversal
- Reranking lowers correct abstention
Plans, code, and graded answers are released.
More from coding & agent
- Microsoft's PrisMem evolves agent memory per-capability, beats baselines by 10.5 points on BEAM-1M — microsoft · 2026-10-07
- No-LLM-agent SRE diagnosis pipeline passes 80/105 cases across 21 fault scenarios in 14.6s median — tianyin_xu · 2026-10-07
- OpenAI staff: chatting with Dot on a 12-hour flight yielded a full workday of output — gabrielchua · 2026-10-07
- Jev-as-a-Judge: New Model Boosts LLM Judge Reliability for Agent Evals — omarsar0 · 2026-10-07
- NVIDIA open-sources OpenShell 0.1.0 to sandbox AI agents without rewriting them — dl_weekly · 2026-10-07
- DeepLearning.AI launches free course on building AI assistants with on-device memory — DeepLearningAI · 2026-10-07