Research: Modifying Just 0.5% of Fine-Tuning Data Can Implant LLM Backdoors

connoraxiotes · x · 2026-08-05

A recent study highlights security risks during the AI model fine-tuning phase. Researchers demonstrated that attackers can implant a subliminal backdoor without any prompt access by modifying only 0.5% of completions. This attack remains effective even if the defender is aware of the behavior and attempts to filter the dataset.

Original post →

More from Safety

Safety channel →