SE-ADD: Self-Evolving Audio Deepfake Detection with Mistake-Driven Supervision

Authors: Rong Wan, Wei Xie, Jiaxi Li, Wenwu Wang, Lu Yin, Yiliao Song, Xilu Wang

Published: 2026-09-30 13:10:57+00:00

AI Summary

This paper introduces SE-ADD, a self-evolving framework for Audio Deepfake Detection (ADD) that adapts Audio Language Models (ALMs) to emerging spoofing attacks. It leverages mistake-driven supervision, built from the ALM's own verdicts and self-generated forensic cues, to iteratively refine the model. SE-ADD demonstrates improved generalization to unseen attacks, significantly reducing Equal Error Rate (EER) on two prominent ALMs.

Abstract

Audio deepfake detection (ADD) must remain effective when new spoofing attacks emerge after deployment. Emerging audio language model (ALM)-based ADD methods are built on predefined supervision from ground-truth labels or verified forensic rationales. However, this paradigm overlooks an ALM's own mistakes, which indicate where targeted supervision is most needed. To this end, we first introduce evolving spoofing environments for ALM-based ADD, where a new attack becomes dominant while previously observed attacks persist. Motivated by the above learning-from-mistakes perspective, we further propose SE-ADD, a self-evolving framework that iteratively adapts an ALM via low-rank adaptation (LoRA) using mistake-driven supervision built from its verdicts and self-generated forensic cues. All training samples receive direct authenticity supervision, while misclassified ones receive additional cue-augmented supervision. As verdicts and cues are regenerated by the updated ALM, the resulting supervision evolves accordingly. Experiments on two ALMs demonstrate the effectiveness of SE-ADD in generalizing to unseen attacks, reducing the equal error rate (EER) from $36.72\\%$ to $7.52\\%$ for Qwen2-Audio and from $19.93\\%$ to $3.97\\%$ for MOSS-Audio.


Key findings
SE-ADD effectively generalizes to unseen attacks, reducing the EER from 36.72% to 7.52% for Qwen2-Audio and from 19.93% to 3.97% for MOSS-Audio. The benefit of forensic-cue supervision is dependent on the inference mode, improving performance under self-hint inference but not always under cue-free inference, and is also backbone-dependent.
Approach
SE-ADD iteratively adapts Audio Language Models (ALMs) using low-rank adaptation (LoRA) within evolving spoofing environments. It generates forensic cues and identifies misclassified samples, providing direct authenticity supervision to all samples and additional cue-augmented supervision specifically to misclassified ones. The updated ALM then regenerates verdicts and cues, leading to an evolving supervision mechanism.
Datasets
ASVspoof 2019 LA train/dev/evaluation partitions
Model(s)
Qwen2-Audio-7B-Instruct, MOSS-Audio-8B-Instruct
Author countries
United Kingdom, China, Australia