Revisiting Cross-Reconstruction for Generalizable Deepfake Detection

Authors: Bingjian Yang, Shilei Zhao, Zheng Wang

Published: 2026-10-01 12:12:46+00:00

AI Summary

The paper revisits cross-reconstruction for image forgery detection, proposing an artifact-oriented disentanglement framework to improve generalization. It argues for preserving artifact diversity, rather than aligning them, and incorporates artifact representations into a masked frequency-aware reconstruction process. This approach helps the model learn transferable forensic representations from diverse manipulation artifacts.

Abstract

Existing image forgery detectors often suffer from generalization to unseen manipulation methods due to the limited ability to capture transferable forensic cues. Recent cross-reconstruction based methods attempt to improve generalization through semantic-artifact disentanglement, but typically align heterogeneous artifacts across generators and exclude artifact representations during reconstruction, which may overlook the inherent diversity and visual cues of manipulation artifacts. In this work, we revisit cross-reconstruction and introduce an artifact-oriented disentanglement framework for robust image forgery detection. We argue that \\textbf{artifact diversity}, i.e., the intrinsic variations of manipulation artifacts introduced by different generation processes, contains complementary forensic cues rather than undesirable domain variations. Instead of enforcing explicit artifact alignment, our framework preserves diverse artifact characteristics through semantically aligned cross-generator reconstruction. Furthermore, we incorporate artifact representations into the reconstruction process and introduce a masked frequency-aware reconstruction strategy to emphasize manipulation-related residuals while reducing semantic interference. This design enables the model to learn transferable forensic representations from diverse artifacts. Extensive experiments on multiple benchmark datasets demonstrate improvements under both cross-dataset and cross-generator evaluation settings. Further analysis and ablation studies validate the effectiveness of artifact diversity preservation and artifact-aware cross-reconstruction.


Key findings
The proposed method achieved state-of-the-art performance in cross-dataset evaluation, with an average AUC of 90.6%, and notably improved video-level detection (e.g., 95.1% AUC on CDF-v2). It also demonstrated competitive performance in cross-generator evaluation, particularly outperforming existing methods on diffusion-based generators like SiT and DiT, highlighting the effectiveness of preserving artifact diversity.
Approach
The authors propose an artifact-oriented disentanglement framework that preserves artifact diversity by using semantically aligned cross-generator reconstruction. They integrate artifact representations into the reconstruction process and employ a masked frequency-aware strategy to emphasize manipulation-related residuals, reducing semantic interference. This enables learning transferable forensic representations for robust deepfake detection.
Datasets
FaceForensics++ (FF++), Celeb-DF-v1/v2, DFDCP, DFDC, DFD, DF40 (VQGAN, StyleGAN-XL, SiT-XL/2, DiT)
Model(s)
ViT-Tiny/16 (visual feature extraction), frozen CLIP ViT-L/14 (classification), AdaIN (Adaptive Instance Normalization)
Author countries
China