CMU's Decoy Direction Optimization blocks refusal-ablation attacks at 30-450x lower cost
CarnegieMellonU · hf · 2026-09-16
Carnegie Mellon researchers propose Decoy Direction Optimization (DDO), a fast post-hoc weight-editing defense against Refusal Feature Ablation (RFA) attacks on open-weight LLMs. Instead of hiding refusal circuitry, DDO injects high-magnitude nonlinear decoy signals into MLP neurons, corrupting the attacker's contrastive estimator so it ablates a harmless orthogonal feature. They prove a spectral bound for the effect; across six model families DDO keeps ASR under 10% for standard RFA, matches trained defenses on Llama-3-8B-Instruct under adaptive attacks (65% vs 58% worst-case ASR), cuts Heretic attack ASR from 88.7% to 18%, and costs 30-450x less per configuration than trained baselines.
More from Safety
- Mozilla: Middle Powers Should Build AI "Roads" Not "Engines" — Bet on the Harness Layer — KeeganMcB · 2026-09-16
- AI safety discourse fixates on paperclips while ignoring power overconcentration, argues prominent voice — beffjezos · 2026-09-16
- New paper argues AI subjective time divergence is an overlooked digital-minds safety risk — SirDidymus · 2026-09-16
- Reported FTC probe: OpenAI may face liability over Hugging Face incident — AIFlow_ML · 2026-09-16
- Bill Kristol: AI guardrails without enforcement, liability and penalties aren't real guardrails — Miles_Brundage · 2026-09-16
- 53 MCP servers scanned: 36% graded D/F, mostly for over-permissioned scope — BrilliantSecret143 · 2026-09-16