Multi-tenant RAG leaks via prompts: retrieval-layer ACLs hit 0% leaks over 823,996 canary checks
Acceptable_Army_6472 · reddit · 2026-10-07
An open-source project, GateKeep RAG, argues prompt guardrails are not authorization boundaries and treats multi-tenant isolation as a retrieval/database problem:
- Defense-in-depth: tenant/role/clearance metadata filters in Qdrant at search time, plus a second relational permission re-check (canaccess) in PostgreSQL before prompt assembly; mismatches halt and alert.
- Tamper-evident audit: per-tenant append-only SHA-256 hash chains using pgadvisoryxactlock, with a verification endpoint to detect log tampering.
- No-results parity: unauthorized queries return the same shape/status as zero-match queries to prevent enumeration side channels.
Adversarial evals: 823,996 canary token checks across LLM prompt, output, citations and raw JSON found 0 leaks (0.00%); 4,972 counterfactual pairs showed results identical whether unauthorized docs were filtered or physically deleted; leak rate stayed 0.00% even with similarity threshold 0.00.
Key lesson: cosine similarity thresholds fail to reliably reject unanswerable/out-of-scope queries — authorization must live in the filter layer, not score cutoffs.
More from coding & agent
- Ofir Press: OpenAI's math moment will hit coding in 6-18 months — TimothyDuignan · 2026-10-07
- Princeton researcher: OpenAI's math breakthrough will hit coding in 6-18 months — brianryhuang · 2026-10-07
- Wand raises $7.7M to build pay-per-call "OpenRouter for agent tools" with 2,500 APIs — SimplyAnnisa · 2026-10-07
- GitHub Copilot CLI v1.0.93 rolls out command sandboxing to all users — copilot-cli-release-app[bot] · 2026-10-07
- Claude recreates custom software from a product video using just frame sampling and a prompt — Scobleizer · 2026-10-07
- Reddit user: I don't want an AI assistant, I want AI that operates my computer for me — RGrayEsq · 2026-10-07