Untraceable DeepFakes via Traceable Fingerprint Elimination

Authors: Jiewei Lai, Lan Zhang, Chen Tang, Pengcheng Sun, Xinming Wang, Yunhao Wang

Published: 2025-08-05 04:27:57+00:00

AI Summary

This paper introduces a novel multiplicative attack designed to create untraceable DeepFakes by fundamentally eliminating traces left by generative models (GMs), a method more robust than existing additive attacks. The proposed universal and black-box attack trains an adversarial model using only real data, effectively circumventing DeepFake attribution models (AMs) and defensive measures. Experimental results demonstrate a high attack success rate against multiple AMs and GMs, highlighting a significant challenge for DeepFake forensics.

Abstract

Recent advancements in DeepFakes attribution technologies have significantly enhanced forensic capabilities, enabling the extraction of traces left by generative models (GMs) in images, making DeepFakes traceable back to their source GMs. Meanwhile, several attacks have attempted to evade attribution models (AMs) for exploring their limitations, calling for more robust AMs. However, existing attacks fail to eliminate GMs' traces, thus can be mitigated by defensive measures. In this paper, we identify that untraceable DeepFakes can be achieved through a multiplicative attack, which can fundamentally eliminate GMs' traces, thereby evading AMs even enhanced with defensive measures. We design a universal and black-box attack method that trains an adversarial model solely using real data, applicable for various GMs and agnostic to AMs. Experimental results demonstrate the outstanding attack capability and universal applicability of our method, achieving an average attack success rate (ASR) of 97.08\\% against 6 advanced AMs on DeepFakes generated by 9 GMs. Even in the presence of defensive mechanisms, our method maintains an ASR exceeding 72.39\\%. Our work underscores the potential challenges posed by multiplicative attacks and highlights the need for more robust AMs.


Key findings
The proposed multiplicative attack achieved an average attack success rate (ASR) of 97.08% against 6 advanced attribution models (AMs) and DeepFakes from 9 generative models (GMs), significantly outperforming state-of-the-art methods. It maintained an ASR exceeding 72.39% even against strong defensive mechanisms like adversarial training, proving the difficulty of defending against such fundamental fingerprint elimination. The method also effectively preserved image fidelity with comparable SSIM and LPIPS metrics.
Approach
The authors propose a universal and black-box multiplicative attack framework, training an adversarial encoder-decoder model solely on real data. This model synthesizes data with artificial fingerprints using sampling and transformation units, then learns to eliminate these fingerprints via a multi-domain loss function (perceptual, spatial, and frequency). The trained model acts as an adversarial matrix to fundamentally eliminate inherent GM fingerprints from DeepFakes, rendering them untraceable while preserving visual fidelity.
Datasets
UNKNOWN
Model(s)
UNKNOWN
Author countries
China