SecOPD Mitigates Adaptive Prompt Injection via On-Policy Distillation

Yibo Peng · hf · 2026-08-27

SecOPD improves defense against adaptive prompt injection attacks by employing on-policy distillation with token-level feedback during fine-tuning. This approach sharply reduces the attack success rate on language models.

Original post →

More from Safety

Safety channel →