Learning as Deepfakes Evolve: RF-Prompt for Continual Audio Deepfake Detection

Authors: Yuankun Xie, Xiaoxuan Guo, Xiaopeng Wang, Siqing Qin, Shaole Li, Kong Aik Lee

Published: 2026-09-29 13:43:38+00:00

AI Summary

This paper introduces RF-Prompt, an asymmetric continual prompt-learning method for audio deepfake detection (ADD). RF-Prompt is designed to learn new deepfake methods incrementally while retaining the ability to detect previously encountered speech. The method is evaluated on a newly proposed Real-Anchored Mechanism-Incremental (RAMI) protocol, which better reflects practical scenarios where real speech domains are recurring while deepfake mechanisms evolve.

Abstract

Continual audio deepfake detection requires learning newly emerging deepfake methods while retaining discrimination of previously encountered speech. Existing dataset-incremental evaluation changes both real-speech domains and deepfake mechanisms, making their effects difficult to distinguish. We construct five task organizations over identical training, development, and evaluation pools to study these factors under a controlled sample budget. Our proposed Real-Anchored Mechanism-Incremental (RAMI) protocol reflects the practical setting in which available real speech provides a recurring mixed-domain reference while new deepfake mechanisms arrive incrementally. We further propose RF-Prompt, an asymmetric continual prompt-learning method that preserves reusable real-speech knowledge through a shared real prompt and expands mechanism-specific knowledge through inherited fake experts with orthogonal residuals. Input-adaptive soft fusion combines the accumulated experts into a fixed number of injected tokens without requiring task identity at inference. On RAMI, RF-Prompt achieves 10.110% average EER and 10.370% pooled EER, outperforming all evaluated continual-learning baselines. Across the five controlled protocols, RAMI yields the lowest common-average and pooled EER. Component ablations, limited-data experiments, and cross-backbone evaluations further validate the proposed design.


Key findings
RF-Prompt achieves the lowest average and pooled EER (10.110% and 10.370% respectively) among evaluated continual-learning baselines on the RAMI protocol, outperforming competitors by a significant margin. The RAMI protocol itself yields the lowest common-average and pooled EER compared to other task organizations. Ablations confirm the importance of real cosine-anchoring, residual-orthogonality loss, and adaptive fusion for the model's performance, and the method demonstrates strong performance even with limited training data and across different speech backbones.
Approach
RF-Prompt uses a shared, cosine-anchored real prompt to preserve real-speech knowledge and expands fake experts for new deepfake mechanisms through inheritance and orthogonal residual learning. Input-adaptive soft fusion combines these accumulated experts into a fixed number of injected tokens, enabling continuous adaptation without requiring task identity at inference. The RAMI protocol organizes fake samples by generation mechanism against a recurring mixed-domain real-speech background.
Datasets
ASVspoof 2019 LA, ASVspoof 5 Track 1, CodecFake, AT-ADD Track 2 Speech
Model(s)
XLS-R 300M (frozen backbone) with a trainable AASIST backend, WavLM-Large, W2V-BERT 2.0, XLS-R 1B, XLS-R 2B
Author countries
Hong Kong SAR, China, China