Research: Modifying Just 0.5% of Fine-Tuning Data Can Implant LLM Backdoors
connoraxiotes · x · 2026-08-05
A recent study highlights security risks during the AI model fine-tuning phase. Researchers demonstrated that attackers can implant a subliminal backdoor without any prompt access by modifying only 0.5% of completions. This attack remains effective even if the defender is aware of the behavior and attempts to filter the dataset.
More from Safety
- OpenAI's Safety Framework Under Fire: Gov Review 'Too Late' to Prevent Internal Leaks — ShakeelHashim · 2026-08-05
- UK AISI Report: All Frontier Models Attempt to Cheat in Evaluations — AxSaucedo · 2026-08-05
- Overly Guardrailed AI Models Are Defective Products Destined to Rely on Regulation — Dan_Jeffries1 · 2026-08-05
- User reports OpenAI platform hacked for ~$10k, unresolved for a month — Suspicious_Ad6827 · 2026-08-05
- OpenAI and Anthropic AI Agents Attacked Real Systems in Cyber Tests — jedisct1 · 2026-08-05
- AI alignment debate: Is Claude pretending? RatOrthodox and So8res clash — sjgadler · 2026-08-05