What Affects the Performance of Fake Audio Detection? Analyzing Factors in a Continual Learning Setting

Authors: Yixuan Xiao, Ngoc Thang Vu

Published: 2026-07-22 11:24:40+00:00

Comment: Accepted to ICASSP25

AI Summary

This study investigates factors influencing fake audio detection (FAD) system performance in a continual learning setup. It analyzes the impact of attacker architectures, training datasets, speaker diversity, and task order on three detection models trained with four different strategies. Key findings indicate that artifacts are often tied to the attackers' training datasets, and both task order and speaker diversity significantly affect detection performance, albeit with varying sensitivity across models and strategies.

Abstract

The increasing sophistication of deepfake audio generation technologies makes it important to develop robust fake audio detection systems that can adapt over time. This study examines how various factors impact the performance of detection systems in a continual learning setting. We focus on factors such as attacker architectures, attackers' training datasets, speaker diversity, and task order. We evaluate the performance of three detection models trained with four different strategies, including direct fine-tuning, one-class classification, random replay, and Learning without Forgetting. Results show that artifacts from the fake audios might arise from the attackers' training datasets, and simply changing attacker architectures does not sufficiently challenge detection systems. Moreover, task order and speaker diversity can significantly influence performance, with varying degrees of sensitivity across different detection models and training strategies. These insights underline the need for careful consideration of these factors when developing robust detection systems.


Key findings
The attacker's training data has a significant impact on detection performance, comparable to the attacker's architecture, a factor often overlooked. Task order can significantly affect continual learning strategies, with LwF being sensitive for smaller models and One-class Softmax becoming unstable with larger models. Speaker diversity generally does not affect most methods, except for One-class Softmax, where high diversity can challenge its underlying premise of compact genuine audio characteristics.
Approach
The authors define tasks by focusing on a single attacker with known architecture and training data, coupled with genuine audios of known diversity, enabling controlled experiments. They evaluate three detection models (Light CNN, ResNet, wav2vec2AASIST) using four training strategies (direct fine-tuning, one-class classification, random replay, Learning without Forgetting) in a continual learning setting to assess factors like attacker's training data vs. architecture, task order, and speaker diversity.
Datasets
ASVSpoof 2019 LA, MLAAD, LibriSpeech
Model(s)
Light CNN (LCNN), ResNet, wav2vec2AASIST
Author countries
Germany