Multi-tenant RAG leaks via prompts: retrieval-layer ACLs hit 0% leaks over 823,996 canary checks

Acceptable_Army_6472 · reddit · 2026-10-07

An open-source project, GateKeep RAG, argues prompt guardrails are not authorization boundaries and treats multi-tenant isolation as a retrieval/database problem:

Adversarial evals: 823,996 canary token checks across LLM prompt, output, citations and raw JSON found 0 leaks (0.00%); 4,972 counterfactual pairs showed results identical whether unauthorized docs were filtered or physically deleted; leak rate stayed 0.00% even with similarity threshold 0.00.

Key lesson: cosine similarity thresholds fail to reliably reject unanswerable/out-of-scope queries — authorization must live in the filter layer, not score cutoffs.

Original post →

More from coding & agent

coding & agent channel →