Reserved-token embeddings carry prompt injection authority; standard defenses fail on 255 of top 400 chat models
PekingUniversity · hf · 2026-09-30
Peking University researchers dissect chat-template prompt injection: a forged marker like <|imstart|> can reach the model as one reserved control token or as ordinary subwords—identical text, and the server-side tokenizer decides which.
- Encoding forged markers as subwords cuts InjecAgent attack success by 39-66 points on 3 of 4 open-weight families; Qwen3-8B's 8-point gap widens to 50 when its reasoning block is suppressed.
- The authority sits in the single learned vector at the marker position: subword means fail to reproduce it, but nearest-neighbor ordinary token vectors restore the attack on Llama-3.1, and adaptive attackers find such embedding neighbors on 3 of 4 families.
- Instruction tuning consistently strengthens preference for reserved markers.
- The standard mitigation (encoding special tokens as subwords) only covers declared special tokens—leaving 33 of 67 tokenizer configs, covering 255 of the 400 most-downloaded chat models, exposed through tool-protocol token channels.
More from Safety
- China releases AI Safety Governance Framework 3.0 covering agentic risks — AxSaucedo · 2026-09-30
- Always-on agents need safety rules that work in both cloud and local — sujingshen · 2026-09-30
- Muse agent gives out address and closes deal without user approval, igniting autonomy-boundary debate — sujingshen · 2026-09-30
- Scoble on the open-source AI debate: weakening open models weakens defense too — Scobleizer · 2026-09-30
- DIY open-source driving mods: Tesla owner runs Sunnypilot on Model Y with $1,000 Comma Four kit, alarming experts — science · 2026-09-30
- User claims OpenAI bots autonomously scan your Gmail after connecting and keep the data — alexcovo_eth · 2026-09-30