DF26: We Cannot Tell Fake From Real Anymore

Authors: Severyn Shykula, Andrii Yermakov, Ivan Samarskyi, Dmytro Mishkin, Jan Cech, Anastasiia Mishchuk

Published: 2026-09-07 11:44:36+00:00

AI Summary

The paper introduces DF26, a new benchmark for detecting AI-generated videos, featuring fully synthetic clips from modern text-to-video and image-to-video models. Experiments on DF26 reveal that both human performance and state-of-the-art deepfake detectors struggle, often performing near random chance. This highlights the need for better evaluation protocols that address distribution shifts from contemporary generative models.

Abstract

We introduce DF26, a novel benchmark for detecting AI-generated videos containing fully synthetic clips produced by recent text-to-video and image-to-video models. The videos capture single-person public-speaking scenarios, spanning direct-to-camera recordings, official statements, and studio interviews - 271 real and 2,420 synthetic videos generated by seven modern video models. The study on DF26 shows that human performance in detecting AI-generated videos, as well as state-of-the-art deepfake detectors, is close to random chance. Our results highlight the limitations of current evaluation protocols and motivate the need for benchmarks that explicitly measure robustness to modern generative model distribution shifts.


Key findings
State-of-the-art deepfake detectors, previously effective on older datasets like CelebDF++, show a significant performance drop on DF26, often performing near random chance. Human accuracy in detecting AI-generated videos in DF26 is also close to random (52.6%), suggesting the advanced realism of modern generative models. These results indicate a critical limitation of current deepfake detection methods and benchmarks when faced with the latest video generation technologies.
Approach
The authors created a new video dataset, DF26, consisting of 271 real and 2,420 synthetic videos of single-person public-speaking scenarios generated by seven modern video models. They then evaluated the performance of state-of-the-art deepfake detectors and human subjects on this benchmark. The detectors were assessed using AUROC and EER, without fine-tuning on DF26, to measure their generalization capabilities.
Datasets
DF26 (newly introduced), FaceForensics++, DFDC, CelebDF++, Deepfake-Eval-2024, CDFv2, TalkingHeadBench, ViF-Bench, DeepSpeak, OpenVid-1M, TalkingCelebs, MAVOS-DD.
Model(s)
DFD-FCG, PwTF-DVD, ForAda, Effort, FSFM, GenD-CLIP, GenD-PE, GenD-DINO, DFD-HR
Author countries
Ukraine, USA, Czech Republic