CMU's Decoy Direction Optimization blocks refusal-ablation attacks at 30-450x lower cost

CarnegieMellonU · hf · 2026-09-16

Carnegie Mellon researchers propose Decoy Direction Optimization (DDO), a fast post-hoc weight-editing defense against Refusal Feature Ablation (RFA) attacks on open-weight LLMs. Instead of hiding refusal circuitry, DDO injects high-magnitude nonlinear decoy signals into MLP neurons, corrupting the attacker's contrastive estimator so it ablates a harmless orthogonal feature. They prove a spectral bound for the effect; across six model families DDO keeps ASR under 10% for standard RFA, matches trained defenses on Llama-3-8B-Instruct under adaptive attacks (65% vs 58% worst-case ASR), cuts Heretic attack ASR from 88.7% to 18%, and costs 30-450x less per configuration than trained baselines.

Original post →

More from Safety

Safety channel →