A Federated Deepfake Speech Detection Method Based on Layer-Wise Center-Guided Weighting Aggregation

Authors: Yingjian Yu, Haiyan Guo, Tianshun Wang, Zirui Ge, Chi Liu

Published: 2026-10-01 07:52:17+00:00

Comment: Accepted and presented at the 17th International Conference on Wireless Communications and Signal Processing (WCSP 2025), Chongqing, China, Oct. 2025

Journal Ref: 2025 17th International Conference on Wireless Communications and Signal Processing (WCSP), 2025

AI Summary

This paper introduces FedDSD, a federated learning approach for deepfake speech detection that allows collaborative model training across decentralized datasets without sharing raw audio, addressing privacy and resource concerns. It utilizes FedProx for local model training to mitigate data heterogeneity and proposes a novel Layer-wise Center-Guided Weighting Aggregation (L-CGWA) strategy for robust global model aggregation. FedDSD achieves comparable performance to centralized co-training while outperforming models trained on individual datasets and demonstrating strong generalization across diverse cross-domain datasets.

Abstract

The advancement of deep learning-based speech synthesis has significantly increased the diversity of deepfake speech, posing threats to voice authentication. While centralized training is effective for deepfake speech detection (DSD), it requires considerable computational resources and raises privacy concerns. To address these issues, we propose a Federated DSD (FedDSD) method that enables collaborative model training across decentralized speech datasets without sharing raw audio. Specifically, each client trains a local model using the FedProx algorithm to mitigate the effects of data heterogeneity and uploads model parameters to a central server. To improve global model aggregation, we further propose a layer-wise center-guided weighting aggregation (L-CGWA) strategy that adjusts each client's contribution per layer based on its distance to a reference center, capturing inter-client and inter-layer discrepancies and enhancing the robustness of model aggregation. Experimental results demonstrate that models trained under the proposed FedDSD method achieve equal error rates (EERs) comparable to those obtained via centralized co-training, while significantly out-performing models trained on individual corpora. Furthermore, the proposed FedDSD method demonstrates robust generalization capabilities across diverse cross-domain datasets.


Key findings
The FedDSD method achieves equal error rates (EERs) comparable to those obtained via centralized co-training, significantly out-performing models trained on individual corpora. The proposed L-CGWA strategy is crucial for this performance, as evidenced by ablation studies, and the method demonstrates robust generalization capabilities across diverse cross-domain datasets, including 21LA, 21DF, and ITW.
Approach
The FedDSD method employs federated learning where clients train local models using FedProx to handle data heterogeneity and then upload parameters to a central server. A novel Layer-wise Center-Guided Weighting Aggregation (L-CGWA) strategy is used for global model aggregation. L-CGWA adjusts each client's layer-wise contribution based on its distance to a reference center, improving robustness by accounting for inter-client and inter-layer discrepancies.
Datasets
ASVspoof2019 LA (19LA), Codecfake, ASVspoof2021 LA (21LA), ASVspoof2021 DF (21DF), In-the-Wild (ITW)
Model(s)
RawBMamba, W2V2-AASIST
Author countries
China