DFALLM: Achieving Generalizable Multitask Deepfake Detection by Optimizing Audio LLM Components

View on arXiv ← Back to list

Authors: Yupei Li, Li Wang, Yuxiang Wang, Lei Wang, Rizhao Cai, Jie Shi, Björn W. Schuller, Zhizheng Wu

Published: 2025-12-09 09:36:38+00:00

AI Summary

This research introduces DFALLM, an optimized Audio Large Language Model (ALLM) framework designed to achieve generalizable and multitask audio deepfake detection. The study investigates the component-level bottleneck in ALLMs, finding that the choice of the audio encoder is crucial for generalization across different spoofing techniques. DFALLM achieves state-of-the-art performance across binary detection, spoof attribution, and spoof localization tasks on multiple public benchmarks.

Abstract

Audio deepfake detection has recently garnered public concern due to its implications for security and reliability. Traditional deep learning methods have been widely applied to this task but often lack generalisability when confronted with newly emerging spoofing techniques and more tasks such as spoof attribution recognition rather than simple binary classification. In principle, Large Language Models (LLMs) are considered to possess the needed generalisation capabilities. However, previous research on Audio LLMs (ALLMs) indicates a generalization bottleneck in audio deepfake detection performance, even when sufficient data is available. Consequently, this study investigates the model architecture and examines the effects of the primary components of ALLMs, namely the audio encoder and the text-based LLM. Our experiments demonstrate that the careful selection and combination of audio encoders and text-based LLMs are crucial for unlocking the deepfake detection potential of ALLMs. We further propose an ALLM structure capable of generalizing deepfake detection abilities to out-of-domain spoofing tests and other deepfake tasks, such as spoof positioning and spoof attribution recognition. Our proposed model architecture achieves state-of-the-art (SOTA) performance across multiple datasets, including ASVSpoof2019, InTheWild, and Demopage, with accuracy reaching up to 95.76% on average, and exhibits competitive capabilities in other deepfake detection tasks such as attribution, and localisation compared to SOTA audio understanding models. Data and codes are provided in supplementary materials.

Key findings

The selection of the audio encoder is the decisive factor for generalizability, with the acoustically-aware Wav2Vec2-BERT outperforming Whisper. The optimal DFALLM configuration achieved a SOTA average detection accuracy of 95.76% across ID and OOD datasets, while also demonstrating superior performance on complex multitask scenarios (attribution and localization) compared to smaller models. Performance analysis suggests that effective deepfake detection requires a high-capacity acoustic encoder but only a lightweight LLM for semantic reasoning.

Approach

DFALLM is a modular ALLM composed of an audio encoder, a linear projection layer, and a textual LLM (fine-tuned with LoRA). The core methodology involves optimizing component selection, specifically prioritizing acoustically-aware audio encoders like Wav2Vec2-BERT over semantic-optimized ones like Whisper. The model uses task-specific textual prompts to perform various deepfake forensic analyses within a unified framework.

Datasets

ASVSpoof2019 (LA), InTheWild (ITW), Demopage, SpoofCeleb, MLAADv6, ReplayDF, DFADD, AISHELL3, ADD2023, GigaSpeech, CNCeleb, PartialSpoof

Model(s)

DFALLM architecture (ALLM); Wav2Vec2-BERT and Whisper (Audio Encoders); Qwen2.5, Qwen3, and Llama (Textual LLMs). The optimal configuration is Wav2Vec2-BERT combined with Qwen2.5-0.5B.

Author countries

UNKNOWN

← Previous