Domain-Adaptive Dual-Gating Mixture of Experts for Generalizable Speech Deepfake Detection

Authors: Siqing Qin, Zhe Li, Kong Aik Lee, Man-Wai Mak

Published: 2026-09-27 16:13:43+00:00

Comment: Accepted by INTERSPEECH 2026

AI Summary

This paper introduces a novel domain-adaptive dual-gating Mixture of Experts (DADGMoE) framework to improve speech deepfake detection (SDD) generalization across unseen attack types and acoustic conditions. The DADGMoE uses a dual-gating mechanism with Sinc-layer-based filters for both low-level acoustic signals and high-level speech representations, guided by domain prototypes for expert routing. This approach significantly reduces the Equal Error Rate (EER) on challenging out-of-dataset benchmarks.

Abstract

Recent advances in speech deepfake detection (SDD) have leveraged the Mixture of Experts (MoE) to enhance generalization capacity. However, existing gating networks often overlook the acoustic and temporal cues of deepfakes. In this work, we propose a novel domain-adaptive dual-gating MoE (DADGMoE) framework for SDD under unseen attack types and acoustic conditions. Our innovative dual-gating mechanism leverages Sinc-layer-based filters to process both low-level acoustic signals (raw waveforms) and high-level speech representations from a large self-supervised learning (SSL) model. It further incorporates domain prototypes to guide expert routing based on implicit deepfake patterns. The lightweight affine experts process the routed inputs. Experiments show that our DADGMoE significantly outperforms the baseline, achieving up to a 40.8% relative EER reduction on challenging out-of-dataset benchmarks. This framework demonstrates superior generalization capabilities and efficient design.


Key findings
The DADGMoE framework achieved significant performance gains, with relative EER reductions of up to 40.8% on out-of-domain datasets like FoR, and 39.3% on ITW. The dual-gating mechanism, combining raw waveform and SSL feature processing, along with domain prototypes, was crucial for improved generalization and expert specialization.
Approach
The DADGMoE framework utilizes a dual-gating mechanism where Sinc-layer-based filters process raw waveforms and high-level SSL features. Domain prototypes guide the routing of these processed inputs to lightweight affine experts. This allows for specialized processing of diverse deepfake patterns, enhancing generalization.
Datasets
ASVspoof 2019 LA (19LA) for training, ASVspoof 2021 Deepfake (21DF), In-the-Wild (ITW), Fake-or-Real (FoR), and ADD 2023 test sets (ADDR1 & ADDR2) for evaluation.
Model(s)
Mixture of Experts (MoE), SincNet, ResNet, XLSR (for SSL features), AASIST (as backend classifier).
Author countries
Hong Kong, China