DBCF: Dual-Branch Complementary Fusion of Foundation Models for Generalized Deepfake Detection

Authors: Fengming Gu, Mingjie He, Zonghui Guo, Jie Zhangb, Shiguang Shan

Published: 2026-09-28 09:26:55+00:00

AI Summary

This paper introduces DBCF, a dual-branch framework for generalized deepfake detection that fuses complementary representations from foundation models. It leverages CLIP for global semantic cues and DINOv3 for fine-grained local structural irregularities. A novel feature fusion module enables parameter-efficient adaptation of these frozen backbones to achieve robust cross-manipulation performance.

Abstract

As image generation and editing technologies have progressed substantially, facial forgeries pose significant challenges to privacy and public safety. Due to limited ability to capture forgery cues, existing small-scale forgery detection models often struggle to generalize across various domains and unseen manipulations. To address this limitation, researchers have turned to large-scale foundation models, which can provide richer representations and better generalization. Nevertheless, relying on a single foundation model alone remains insufficient for effective forgery detection. While models like CLIP offer robust global semantic cues, they lack the capacity to capture detailed local facial features. In contrast, DINO excels at capturing local structural features of faces, but provides weaker global semantic context. To fully utilize the synergies among multiple foundation models, we propose a hierarchical multi-granular framework that integrates complementary pretrained representations. Specifically, a Global Context Branch (GCB) based on CLIP captures holistic semantic cues, while a Fine-grained Cue Branch (FCB) built on DINOv3 captures localized structural irregularities. In addition, we design a feature fusion module that enables parameter-efficient adaptation of the frozen foundation backbones by adaptively extracting and integrating complementary features from the two models. By jointly leveraging global context and fine-grained cues, our method learns more comprehensive forgery representations and achieves strong cross-manipulation performance. Extensive experiments on multiple benchmarks demonstrate the benefit of the proposed design, particularly under cross-dataset and cross-manipulation settings.


Key findings
DBCF achieves strong generalization performance in cross-dataset and cross-manipulation settings, outperforming state-of-the-art methods on several benchmarks, particularly on challenging datasets like DFDC and DFDCP. The dual-branch approach, combining global semantic and fine-grained local cues, significantly enhances robustness against unseen manipulation types and image perturbations. The ablation studies confirm that the synergistic integration of both branches and the multi-granular feature aggregation are crucial for the improved performance, showing a reasonable trade-off between performance gain and computational overhead.
Approach
The proposed DBCF framework employs two main branches: a Global Context Branch (GCB) utilizing a frozen CLIP visual backbone to capture holistic semantic cues, and a Fine-grained Cue Branch (FCB) built on a frozen DINOv3 backbone to extract localized structural irregularities. An Adaptive Feature Learner (AFL) extracts task-adaptive multi-scale spatial priors, which are then progressively refined and integrated with the GCB and FCB features through Cross-Feature Interaction (CFI) modules. Finally, a multi-scale decoder aggregates these hierarchical features for real/fake prediction.
Datasets
FaceForensics++ (FF++), Celeb-DF v2, DFDC, DFDCP, DF40 (for cross-manipulation evaluation)
Model(s)
CLIP-ViT-L/14-336 (for GCB), DINOv3-ViT-L/16 (for FCB)
Author countries
China