GLAD: Global-Local Adaptive Detector for Robust Speech Deepfake Detection
Authors: Zelin Zhao, Guanjie Huang, Danny Hin Kwok Tsang, Li Liu
Published: 2026-09-28 15:29:03+00:00
AI Summary
This paper introduces GLAD, a Global-Local Adaptive Detector designed to overcome the limitations of current Speech Deepfake Detection (SDD) models, particularly their poor generalization to unseen domains and inability to capture fine-grained signal artifacts. GLAD employs a Hierarchical Global-Local (HGL) backbone to fuse global linguistic and acoustic features with local signal details, a Hierarchical Adaptive Gating (HAG) mechanism for dynamic layer-wise focus, and SaniBoost for robust data augmentation. Extensive experiments demonstrate GLAD's superior performance, especially in out-of-distribution scenarios.
Abstract
Recent advances in AI-based speech synthesis have enabled highly realistic speech, increasing the importance of speech deepfake detection (SDD) in preventing misuse. While mainstream Self-Supervised Learning (SSL)-based detectors achieve strong performance, they suffer from poor generalization to unseen domains and often overlook fine-grained signal artifacts due to a bias towards global semantic consistency. In this paper, we conduct the first detailed empirical and visual analysis to validate these limitations explicitly. Our investigation reveals two critical architectural vulnerabilities: (1) a systemic failure to capture localized spoofing traces, and (2) a severe lack of adaptability to domain-driven shifts in SSL layer importance, rendering static aggregation strategies prone to overfitting. To address these vulnerabilities, we propose the Global-Local Adaptive Detector (GLAD). Specifically, to capture localized forgeries, GLAD employs a Hierarchical Global-Local (HGL) backbone that explicitly bridges the granularity gap by fusing global linguistic and acoustic features with fine-grained local signal details. To counter layer importance shifts in out-of-distribution (OOD) scenarios, we introduce a Hierarchical Adaptive Gating (HAG) mechanism that dynamically recalibrates layer-wise focus in a sample-specific manner. Finally, to address shortcut learning induced by environmental biases, we introduce SaniBoost, a composite data augmentation strategy for robust signal standardization and noise sanitization. Extensive experiments demonstrate that GLAD significantly outperforms state-of-the-art methods, particularly on unseen domain cases.The code will be released upon publication.