Graph memory loses to flat vector baseline on LongMemEval: F1 0.42 vs 0.47
CShorten30 · x · 2026-09-06
A controlled study testing whether graph memory beats flat retrieval for long-term agents found the opposite: researchers extracted each conversational turn into typed nodes and attributed edges, answered from two-hop subgraphs, and held the candidate-generation budget at five retrieval roots. On LongMemEval the graph scored token F1 0.42 vs 0.47 for the flat baseline, with a paired bootstrap over 500 questions showing a -0.050 gap (95% CI -0.085 to -0.016). The damage concentrated on questions requiring recall of a specific prior assistant turn, where judged correctness fell from 0.911 to 0.607—splitting turns into entities discards the surface form those questions need.
More from coding & agent
- Designing a multi-agent due diligence app to audit online gurus and e-commerce brands — AgentVN · 2026-09-06
- AutoHedge: open-source swarm-agent framework builds an autonomous hedge fund in minutes — The-Swarm-Corporation · 2026-09-06
- OpenAI open-sources official Skills Catalog for Codex, 25k stars on GitHub — openai · 2026-09-06
- Cryptographer Matthew Green uses coding agents to teach: the code isn't the goal, the student is — matthew_d_green · 2026-09-06
- Codex 'godmode' config: 1M context window, low reasoning effort and a 400K auto-compact limit — EXM7777 · 2026-09-06
- Cryptography professor Matthew Green redesigns course for AI: agents allowed, more exams — matthew_d_green · 2026-09-06