Deepfake Detectors Decay to 76% Accuracy on 2024 Generators; Researchers Propose Calibrated Authentication
MBZUAI · hf · 2026-10-06
MBZUAI researchers evaluate twenty deepfake detectors against ten generators released over the last four years and identify a fundamental flaw in detection paradigms.
- Detector decay: Accuracy falls from near-perfect 99.5% to 76% over time; adversarial perturbations drop every baseline below 2% accuracy, even inverting assigned labels.
- Root ambiguity: Generators can reproduce authentic content exactly (e.g., via memorization), so content alone can't reveal provenance—generator-produced content must admit a faithful reconstruction by that same generator, making authenticity plausibly deniable.
- New paradigm: They propose calibrated prediction of plausible deniability, certifying at most 1% of generated content wrongly—an operating point where most baselines, including the strongest (93% accuracy), reach near-zero recall—and preserving this bound against bounded-perturbation adaptive attacks.
- Eroding verifiability: Of 3,000 Reddit images, 1,116 resist reproduction by a 2022 generator but only 55-79 resist 2024 generators.
More from Safety
- Anti-AI protesters turn to direct action, disrupting Nvidia dinner with 'pull the plug' banner — nordicinst · 2026-10-06
- 38% of AI Agent Container Escapes Needed No Kernel 0-Days: Analysis of 109 Incidents — doletskyisergey · 2026-10-06
- Freelancer finds ChatGPT voice chats leaking into transcription gig work — MilagrosMiceli · 2026-10-06
- Agent safety debate: capability sets damage size, but 'orphanhood' decides accountability — mariotelfig · 2026-10-06
- Closed-Loop Attack Injects Bias into Diffusion LLMs in 40 Minutes on One GPU — MBZUAI · 2026-10-06
- Neel Nanda: Anti-Safety PACs Outspend Pro-Safety Groups on AI Policy — NeelNanda5 · 2026-10-06