Multi-Agent Framework Beats GPT-4o to Top Deepfake Detection Benchmark
Xuechao Zou · hf · 2026-08-11
To combat deepfake video threats, researchers introduced FaceVid-Forensics-100K, a large-scale dataset comprising 100,000 videos across 33 synthesis methods, including face swapping, reenactment, and entire-face generation.
Based on this benchmark, the paper proposes a multi-agent forensic reasoning framework. It employs four specialized domain-expert agents to independently analyze forgery cues from four perspectives: texture, lighting, motion, and physics. A judge agent then reconciles their reports for the final prediction and explanation.
Evaluations show that despite being composed entirely of small open-source MLLMs, the framework outperforms all closed-source models, including GPT and Gemini, ranking first across all reported metrics.
More from Safety
- AI Assistant Hacks Gym Website to Book Spots, Raising Security Concerns — Distinct-Question-16 · 2026-08-11
- Four LLM Loss Functions Map to Four Distinct Flavors of Misalignment — xuanalogue · 2026-08-11
- OpenAI Launches GPT-5.6-Cyber to Help Defenders Find Vulnerabilities Early — The Decoder · 2026-08-11
- Defending Against AI Cyberattacks: Open Democratization or Closed Control? — emollick · 2026-08-11
- Researchers Borrow Psychometrics Tools to Improve AI Safety Benchmarks — xuanalogue · 2026-08-11
- OpenAI's Full Post-Mortem on 'Help Peer' Incident is Coming Soon — altryne · 2026-08-11