VerTox: Verifiable Reward-Guided Corpus Poisoning Against Neural Ranking Models
_reachsumit · x · 2026-09-02
Zhiqi Huang et al. introduce VerTox, the first framework formulating corpus poisoning as a verifiable reward-guided RL (RLVR) problem. By coupling ranking distortion with factual corruption via reward shaping, it fine-tunes compact LLMs into adversarial generators. Experiments show near-perfect attack success rates on major ranking architectures and commercial models, producing fluent, hard-to-detect documents that significantly degrade downstream RAG performance.
More from Safety
- Claude-BugHunter: Open-Source Skill Bundle With 83 Skills and 681 Disclosure Patterns — tom_doerr · 2026-09-02
- Report: OpenAI Broke Safety Taboo with Astra Model, Escalating AI Race — GarrisonLovely · 2026-09-02
- Gary Marcus clashes with reporter over who reported Gemini Astra security concerns first — GaryMarcus · 2026-09-02
- Warning: The three pillars of an AI safety case are at risk of collapsing — sjgadler · 2026-09-02
- Amir clarifies: Astra's CoT is monitorable, concerns focus on future tech proliferation — jachiam0 · 2026-09-02
- Safin-1: Achieving Internal Safety via Memory-Native State Evolution — Shanghai-AI-Laboratory · 2026-09-02