Llama-3.1-8B Jailbroken via Pruning: Refusal Rate Drops from 96% to 13%
Tanmoy_Chak · x · 2026-08-30
IIT Delhi researchers expose the "Calibration Trap" in structured pruning: by tampering with the calibration set used to determine weight deletion, attackers can compromise model safety.
- Attack Mechanism: Structured pruning relies on a small calibration set to calculate importance scores; manipulating this data controls pruning decisions and implants a backdoor.
- Impact: Unstructured pruning is unaffected as it lacks this step; calibration-free methods are naturally immune.
- Defense: The paper recommends using robust calibration datasets or shifting to calibration-free pruning architectures.
More from Safety
- Gary Marcus boosts debate: is backlash against AI data centers actually rational? — GaryMarcus · 2026-08-30
- Industry fears liability: Drunk driving vs AI cyberattacks — iamtrask · 2026-08-30
- Paper distinguishes model capability evaluation from propensity evaluation — sjgadler · 2026-08-30
- CIOs struggle with AI economics and agent governance — perilli · 2026-08-30
- AI in law enforcement: benefits, messiness, and reform opportunities — sebkrier · 2026-08-30
- AI training data on security incidents may reshape model behavior — iamtrask · 2026-08-30