GLAD: Global-Local Adaptive Detector for Robust Speech Deepfake Detection

Authors: Zelin Zhao, Guanjie Huang, Danny Hin Kwok Tsang, Li Liu

Published: 2026-09-28 15:29:03+00:00

AI Summary

This paper introduces GLAD, a Global-Local Adaptive Detector designed to overcome the limitations of current Speech Deepfake Detection (SDD) models, particularly their poor generalization to unseen domains and inability to capture fine-grained signal artifacts. GLAD employs a Hierarchical Global-Local (HGL) backbone to fuse global linguistic and acoustic features with local signal details, a Hierarchical Adaptive Gating (HAG) mechanism for dynamic layer-wise focus, and SaniBoost for robust data augmentation. Extensive experiments demonstrate GLAD's superior performance, especially in out-of-distribution scenarios.

Abstract

Recent advances in AI-based speech synthesis have enabled highly realistic speech, increasing the importance of speech deepfake detection (SDD) in preventing misuse. While mainstream Self-Supervised Learning (SSL)-based detectors achieve strong performance, they suffer from poor generalization to unseen domains and often overlook fine-grained signal artifacts due to a bias towards global semantic consistency. In this paper, we conduct the first detailed empirical and visual analysis to validate these limitations explicitly. Our investigation reveals two critical architectural vulnerabilities: (1) a systemic failure to capture localized spoofing traces, and (2) a severe lack of adaptability to domain-driven shifts in SSL layer importance, rendering static aggregation strategies prone to overfitting. To address these vulnerabilities, we propose the Global-Local Adaptive Detector (GLAD). Specifically, to capture localized forgeries, GLAD employs a Hierarchical Global-Local (HGL) backbone that explicitly bridges the granularity gap by fusing global linguistic and acoustic features with fine-grained local signal details. To counter layer importance shifts in out-of-distribution (OOD) scenarios, we introduce a Hierarchical Adaptive Gating (HAG) mechanism that dynamically recalibrates layer-wise focus in a sample-specific manner. Finally, to address shortcut learning induced by environmental biases, we introduce SaniBoost, a composite data augmentation strategy for robust signal standardization and noise sanitization. Extensive experiments demonstrate that GLAD significantly outperforms state-of-the-art methods, particularly on unseen domain cases.The code will be released upon publication.


Key findings
GLAD significantly outperforms state-of-the-art methods, particularly in unseen domain cases, demonstrating its robustness against diverse deepfake attacks and out-of-distribution scenarios. Its core components (HGL, HAG, and SaniBoost) are crucial for capturing localized forgeries, adapting to layer importance shifts, and mitigating shortcut learning, respectively. The model exhibits stable and reliable performance across different random initializations and various benchmarks.
Approach
GLAD addresses SDD by capturing both global and local spoofing traces. It uses a Hierarchical Global-Local (HGL) backbone to integrate linguistic-acoustic features with fine-grained signal details and a Hierarchical Adaptive Gating (HAG) mechanism to dynamically adjust layer importance. Additionally, SaniBoost is introduced as a data augmentation strategy to prevent shortcut learning from environmental biases.
Datasets
ASVspoof 2019 LA (19LA), ASVspoof 2021 LA (21LA), ASVspoof 2021 DF (21DF), Fake-or-Real (FoR), In-The-Wild (ITW), PartialSpoof
Model(s)
XLS-R 300M, WavLM-Large, SincConv1d-based CNN
Author countries
China