Multi-Agent Framework Beats GPT-4o to Top Deepfake Detection Benchmark

Xuechao Zou · hf · 2026-08-11

To combat deepfake video threats, researchers introduced FaceVid-Forensics-100K, a large-scale dataset comprising 100,000 videos across 33 synthesis methods, including face swapping, reenactment, and entire-face generation.

Based on this benchmark, the paper proposes a multi-agent forensic reasoning framework. It employs four specialized domain-expert agents to independently analyze forgery cues from four perspectives: texture, lighting, motion, and physics. A judge agent then reconciles their reports for the final prediction and explanation.

Evaluations show that despite being composed entirely of small open-source MLLMs, the framework outperforms all closed-source models, including GPT and Gemini, ranking first across all reported metrics.

Original post →

More from Safety

Safety channel →