Fit for Purpose? Deepfake Detection in the Real World

Authors: Guangyu Lin, Li Lin, Christina P. Walker, Daniel S. Schiff, Shu Hu

Published: 2025-10-18 16:00:10+00:00

AI Summary

This paper introduces the Political Deepfakes Incident Database (PDID), the first systematic benchmark of real-world political deepfakes circulating on social media. It conducts a comprehensive evaluation of state-of-the-art deepfake detectors from academia, government, and industry on this new dataset. The study reveals that most existing detection models struggle to generalize to authentic political deepfakes and are vulnerable to simple manipulations, particularly in the video domain.

Abstract

The rapid proliferation of AI-generated content, driven by advances in generative adversarial networks, diffusion models, and multimodal large language models, has made the creation and dissemination of synthetic media effortless, heightening the risks of misinformation, particularly political deepfakes that distort truth and undermine trust in political institutions. In turn, governments, research institutions, and industry have strongly promoted deepfake detection initiatives as solutions. Yet, most existing models are trained and validated on synthetic, laboratory-controlled datasets, limiting their generalizability to the kinds of real-world political deepfakes circulating on social platforms that affect the public. In this work, we introduce the first systematic benchmark based on the Political Deepfakes Incident Database, a curated collection of real-world political deepfakes shared on social media since 2018. Our study includes a systematic evaluation of state-of-the-art deepfake detectors across academia, government, and industry. We find that the detectors from academia and government perform relatively poorly. While paid detection tools achieve relatively higher performance than free-access models, all evaluated detectors struggle to generalize effectively to authentic political deepfakes, and are vulnerable to simple manipulations, especially in the video domain. Results urge the need for politically contextualized deepfake detection frameworks to better safeguard the public in real-world settings.


Key findings
Detectors from academia and government perform relatively poorly on real-world political deepfakes, while paid commercial tools and LVLMs achieve higher performance but still struggle with generalization and often exhibit high false acceptance rates. All evaluated detectors find political video deepfake detection particularly challenging and are vulnerable to simple manipulations, highlighting the urgent need for politically contextualized detection frameworks.
Approach
The authors curated the Political Deepfakes Incident Database (PDID) of real-world political deepfakes (images and videos) with human-verified labels sourced from social media. They then systematically benchmarked a wide array of deepfake detectors, including academic white-box models, government-developed tools, commercial black-box solutions, and large vision-language models (LVLMs), evaluating their performance and robustness on the PDID dataset.
Datasets
Political Deepfakes Incident Database (PDID)
Model(s)
Xception, EfficientNet-B4, ViT-B/16, F3Net, SPSL, SRM, UCF, UnivFD, CORE, DAW-FDD, DAG-FDD, PG-FDD, CNNDetection, KitwareDetector, GANattribution, Incode, Is It AI, Winston, BrandWell, AI or Not, Illuminarty, Reality Defender, Hive Moderation, Deepfake Detector, ChatGPT(GPT-5), Qwen-VL-Max, Claude-Sonnet-4.5, Gemini-2.5-pro, LLaVA-v1.5-7B, LLaVA-v1.5-13B, LLaVA-v1.5-13B-XTuner, DeepSeek-VL-7B, CogVLM-Chat, Monkey-Chat
Author countries
USA