Dev benchmark: keyword search hits 1.000 recall while vector RAG drops to 0.839
Initial_Orange2985 · reddit · 2026-09-29
A developer building a memory layer for an assistant published a candid retrieval benchmark comparing it against vector RAG and keyword search — with counterintuitive results.
Setup:
- 180 questions per size, histories of 104/208/416 facts, k=6, zero LLM calls; measuring whether the needed fact is actually injected
- Vector RAG (FAISS, same product embeddings) degrades from 0.889 to 0.839 as history grows
- Keyword top-k and the memory layer both score 1.000 everywhere
Findings:
- The author initially thought keyword search scored only 0.672 — turned out they weren't splitting underscores in relation names, so keywords couldn't match. After fixing, the gap vanished: a humbling lesson about baselines
- The memory layer's real edge over keywords is size: 134–177 chars injected per query vs 380 — same recall with 2–3x less context; and it admits "nothing in memory" instead of stuffing 6 half-related lines
- Limitations: synthetic French corpus, literal keys, retrieval-only
- Next benchmark planned: a stated fact contradicted 8 times by an imported doc, where duplicates can fill all k=6 slots — asking the community about dedup and source-tagging strategies
More from coding & agent
- OpenAI DevDay live thread: 21 updates, headline release likely the always-on agent — kimmonismus · 2026-09-30
- Apodex 1.1 launches: agents can re-plan mid-task without restarting research — SucceededMind · 2026-09-30
- Neuroscience Researcher Uses Claude and GPT as Rival Research Assistants — Now Unsure Where to Publish — neuroecology · 2026-09-30
- Structure-only AI audit: Jev predicts outcomes of 2,029 real calls at AUC 0.78 for $3 — alexcovo_eth · 2026-09-30
- Replicas V3 launches: run Claude Code and Codex in cloud VMs with BYO keys — KlausCodes · 2026-09-30
- Agno adds first-class MCP publishing for agent components — pritisinghhhh · 2026-09-30