DF26 Benchmark: Humans and SOTA Detectors Spot AI Videos Near Random Chance

Justgototheeffinmoon · reddit · 2026-09-11

Researchers introduce DF26, a new benchmark for detecting AI-generated videos: 271 real clips plus 2,420 synthetic videos from seven modern text-to-video and image-to-video models, all featuring single-person public-speaking scenarios.

Key findings:

The authors argue current evaluation protocols are inadequate and call for benchmarks that explicitly measure robustness to distribution shifts from modern generative models. Bottom line: we can no longer tell fake from real.

Related event: DF26 benchmark: humans and detectors near random on AI speech videos(3 posts)→

Original post →

More from Safety

Safety channel →