InceptionRAG: Dormant-Passage Poisoning Attack Hits 80%+ Success Against RAG Defenses
chaumian · x · 2026-09-16
A new arXiv paper, InceptionRAG, reframes RAG corpus poisoning: instead of embedding an explicit malicious payload in one document, it fragments the attack into a chain of individually harmless "dormant passages" that slip past current mitigations. When retrieved together, they push LLMs to self-deduce target misinformation via multi-hop reasoning. A zeroth-order suffix optimization (ZOSO) method automates authoritative suffix generation for black-box settings. Across 3 datasets and 3 LLMs, the attack exceeds an 80% success rate even under rigorous defenses.
More from Safety
- Anthropic and OpenAI spend only ~1% of capex on AI safety, analysis shows — AlexTensor · 2026-09-16
- 53 MCP servers scanned: 36% graded D/F, mostly for over-permissioned scope — BrilliantSecret143 · 2026-09-16
- CMU's Decoy Direction Optimization blocks refusal-ablation attacks at 30-450x lower cost — CarnegieMellonU · 2026-09-16
- US voters oppose AI data centers 57%-71%, while DOJ backs OpenAI's fair-use defense vs NYT — emmanuelvivier · 2026-09-16
- Common Sense rates Perplexity an 'unacceptable risk' for children, below Google — emmanuelvivier · 2026-09-16
- Universal Music sues DistroKid over 'AI slop factory', citing 1,000 infringed works — emmanuelvivier · 2026-09-16