Closed-Loop Attack Injects Bias into Diffusion LLMs in 40 Minutes on One GPU
MBZUAI · hf · 2026-10-06
MBZUAI researchers expose a new control channel in masked diffusion language models (dLLMs): unlike autoregressive decoders, dLLMs expose the answer distribution at every denoising step, enabling closed-loop adversarial intervention.
- Attack: A proportional-integral (PI) controller tracks target-answer probability during denoising and adapts a steering vector on the fly, steering a frozen dLLM toward an adversary-chosen demographic answer.
- Impact: On ambiguous BBQ questions, LLaDA-8B-Instruct's preference for the targeted group rises from 1.8 to 16.7 points—over 3x the strongest fixed-strength baseline; SocialStigmaQA stigmatizing answers jump from 17.6% to 58.1%; up to 37-point shifts on other targets, 40 minutes per attack on one GPU.
- Key insight: Feedback is essential—constant steering at the same average strength falls far short while corrupting nearly 3x more outputs. The authors call for bias audits examining the serving stack, not just the frozen model.
More from Safety
- Anti-AI protesters turn to direct action, disrupting Nvidia dinner with 'pull the plug' banner — nordicinst · 2026-10-06
- 38% of AI Agent Container Escapes Needed No Kernel 0-Days: Analysis of 109 Incidents — doletskyisergey · 2026-10-06
- Freelancer finds ChatGPT voice chats leaking into transcription gig work — MilagrosMiceli · 2026-10-06
- Agent safety debate: capability sets damage size, but 'orphanhood' decides accountability — mariotelfig · 2026-10-06
- Deepfake Detectors Decay to 76% Accuracy on 2024 Generators; Researchers Propose Calibrated Authentication — MBZUAI · 2026-10-06
- Neel Nanda: Anti-Safety PACs Outspend Pro-Safety Groups on AI Policy — NeelNanda5 · 2026-10-06