A War Beyond Deepfake: Benchmarking Facial Counterfeits and Countermeasures

Authors: Minh Tam Pham, Thanh Trung Huynh, Van Vinh Tong, Thanh Tam Nguyen, Thanh Thi Nguyen, Hongzhi Yin, Quoc Viet Hung Nguyen

Published: 2021-11-25 05:01:08+00:00

AI Summary

This paper presents a comprehensive benchmark for visual facial forgery and forensics, integrating state-of-the-art deepfake generators and detectors within an independent framework. It measures and analyzes their performance under various criteria and adversarial conditions to provide insights and guidelines for the ongoing struggle between forgery techniques and detection countermeasures. The study also performs an exhaustive analysis of benchmarking results to determine characteristics and serve as a comparative reference.

Abstract

In recent years, visual forgery has reached a level of sophistication that humans cannot identify fraud, which poses a significant threat to information security. A wide range of malicious applications have emerged, such as fake news, defamation or blackmailing of celebrities, impersonation of politicians in political warfare, and the spreading of rumours to attract views. As a result, a rich body of visual forensic techniques has been proposed in an attempt to stop this dangerous trend. In this paper, we present a benchmark that provides in-depth insights into visual forgery and visual forensics, using a comprehensive and empirical approach. More specifically, we develop an independent framework that integrates state-of-the-arts counterfeit generators and detectors, and measure the performance of these techniques using various criteria. We also perform an exhaustive analysis of the benchmarking results, to determine the characteristics of the methods that serve as a comparative reference in this never-ending war between measures and countermeasures.


Key findings
State-of-the-art visual forensic techniques generally lack generalization to unseen forgery methods, though expert-defined features can mitigate this limitation. Deep learning models like XceptionNet and Capsule perform best in ideal conditions but show reduced effectiveness with noise or low resolution, while GAN-fingerprint proves more robust in such extreme input conditions. Illumination factors such as brightness and contrast significantly influence detection performance, and the localization of suspicious regions by forensic models varies considerably depending on the forgery type.
Approach
The authors developed an independent, extensible framework that integrates various state-of-the-art visual forgery generators (e.g., DeepFake, StarGAN) and forensic detectors (e.g., XceptionNet, Capsule). They conduct a comprehensive dual benchmarking study, evaluating the performance of these techniques under various adversarial conditions such as brightness, contrast, noise, resolution, missing information, and compression. The framework allows for in-depth analysis of generalization abilities and feature overlapping among forgery types.
Datasets
CelebA-HQ, DFDC, FaceForensic++, DeepFake-in-the-wild, Celeb-DF, UADFV, DF-TIMIT
Model(s)
UNKNOWN
Author countries
Vietnam, Australia, Germany