Dev Benchmarks Memory Layer vs Vector RAG and Keyword Search: Keyword Tied, Vector RAG Fell

Initial_Orange2985 · reddit · 2026-09-29

A developer benchmarked his assistant's memory layer on one question: is the fact the model needs actually in what gets injected? Setup: 180 questions per size, 104/208/416 facts in history, same corpus and normalization, k=6 for every arm, zero LLM calls.

Results: vector RAG (FAISS, same embeddings as his product) scored 0.889, 0.850, 0.839, dropping as history grows; keyword top-k scored 1.000 at all sizes; his memory layer also scored 1.000. So keyword search ties him on recall.

He notes his first version had keyword at 0.672 and he thought he was crushing it, until he realized he wasn't splitting underscores in relation names so it couldn't match. Fixing it erased the gap — a lesson about baselines.

Where memory wins is size: 134–177 chars injected per question vs 380 for keywords, same recall with 2–3x less context. And when it has nothing, it says "nothing in memory on this" instead of shoving in six half-related lines.

He admits limits: synthetic corpus, French, literal keys, retrieval only, not end-to-end. Next he's building a nastier test where a user states a fact and an imported doc repeats the opposite eight times, potentially filling all k=6 slots, and asks how others handle it.

Related event: Benchmarks Show Keyword Search Matches or Beats Vector RAG for Agent Memory Retrieval(2 posts)→

Original post →

More from coding & agent

coding & agent channel →