DF26 Benchmark: Humans and SOTA Detectors Spot AI Videos Near Random Chance
Justgototheeffinmoon · reddit · 2026-09-11
Researchers introduce DF26, a new benchmark for detecting AI-generated videos: 271 real clips plus 2,420 synthetic videos from seven modern text-to-video and image-to-video models, all featuring single-person public-speaking scenarios.
Key findings:
- Human performance at spotting AI videos is close to random chance
- State-of-the-art deepfake detectors perform near random as well
The authors argue current evaluation protocols are inadequate and call for benchmarks that explicitly measure robustness to distribution shifts from modern generative models. Bottom line: we can no longer tell fake from real.
Related event: DF26 benchmark: humans and detectors near random on AI speech videos(3 posts)→
More from Safety
- AI safety insider: cheap extinction-risk talk is destroying the public's imagination of AI — Dr_Atoosa · 2026-09-11
- PuzzleMask: Plain-Prose Attack Bypasses All 4 Tested LLM Gatekeepers at 100% — TechNadu · 2026-09-11
- Anthropic Says It Blocked Attempts to Use AI for Bioweapons Development — KoseteBamse · 2026-09-11
- Never hardcode AI API keys: attackers scan app binaries, GitHub and Docker — eyishazyer · 2026-09-11
- Beware hotel Wi-Fi popups: DNS hijacking used to deliver malware — eyishazyer · 2026-09-11
- First $1B AI-agent breach may look like software working as designed, says Enigma CTO — TechNadu · 2026-09-11