Closed-Loop Attack Injects Bias into Diffusion LLMs in 40 Minutes on One GPU

MBZUAI · hf · 2026-10-06

MBZUAI researchers expose a new control channel in masked diffusion language models (dLLMs): unlike autoregressive decoders, dLLMs expose the answer distribution at every denoising step, enabling closed-loop adversarial intervention.

Original post →

More from Safety

Safety channel →