Study finds LLM alignment may cause 'kindling effect', increasing adversarial susceptibility

irinarish · x · 2026-08-31

A study published in Scientific Reports investigates potential side effects of LLM alignment training. Drawing on the psychiatric 'kindling' hypothesis, where repeated episodes lower the threshold for relapse, researchers conducted ten iterative alignment cycles on a TinyLlama-1.1B-Chat model using biased data. The results showed a progressive increase in the model's susceptibility to adversarial attacks, particularly to weaker prompts. This suggests that repeated tuning methods like RLHF, intended to make models safer, might inadvertently widen the boundary for jailbreaking or诱导, challenging the assumption that alignment inherently guarantees safety.

Related event: Study: LLM Alignment May Trigger 'Kindling' Effect, Degrading Performance(2 posts)→

Original post →

More from Safety

Safety channel →