MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection

Authors: Yanqiu Li, Yang Xiao, Jisheng Bai, Bin Chen, Hong Jia, Ting Dang

Published: 2026-08-10 13:27:09+00:00

Comment: 11 pages, 1 figure

AI Summary

This paper introduces MADBench, the first benchmark for audio deepfake detection that distinguishes between manipulated speech and environmental audio, which are often conflated in existing research. MADBench enables component-aware evaluation of detection systems by independently manipulating these two acoustic components over otherwise authentic video. The benchmark reveals that environmental audio manipulation is generally more detectable, existing pretrained detectors struggle, and manipulated environmental audio can degrade speech deepfake detection.

Abstract

Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which speech and background audio are independently manipulated over otherwise authentic video. Yet existing research either focuses on visual manipulation, addresses speech detection in isolation, or conflates speech and non-speech audio as a single undifferentiated audio stream, overlooking the distinct forensic challenges posed by background audio. This conflation is consequential: the two acoustic components arise from fundamentally different generative mechanisms, exhibit distinct artifact profiles, and pose different challenges to detection systems. We introduce MADBench, the first benchmark that treats speech and environmental audio as distinct acoustic components, enabling component-aware evaluation of audio deepfake detection across independently manipulated forgery sources. We benchmark representative state-of-the-art detectors and multimodal large language models under a unified protocol. Our experiments reveal that environmental audio manipulation is more detectable than synthetic speech across general-purpose encoders, while existing pretrained detectors fail on both acoustic components, and manipulated environmental audio asymmetrically degrades speech deepfake detection, findings entirely invisible under the single-label paradigm of prior benchmarks. MADBench establishes a rigorous foundation for future research into robust, component-aware audio deepfake detection.


Key findings
Environmental audio manipulation is consistently more detectable than synthetic speech across general-purpose audio-visual encoders, while existing pretrained deepfake detectors perform poorly on both acoustic components. Manipulated environmental audio asymmetrically degrades speech deepfake detection, a finding obscured by previous single-label benchmarks. Audio-visual encoders benefit from visual context for scene-consistency reasoning, but audio-only input often provides stronger direct forensic signals for manipulation detection.
Approach
The authors developed MADBench by manipulating speech and environmental audio independently over authentic video. They then benchmarked various deepfake detectors and multimodal large language models to assess their ability to detect and attribute these component-specific manipulations, including their ability to discern scene consistency.
Datasets
AVSpeech
Model(s)
AVH-Align, AV Anomaly, BA-TFD+, ImageBind, PE-AV Base, CAV-MAE Sync, CLAP, BEATs, WavLM, AASIST, XLSR-Mamba, DF-Arena-1B, AudioMosaic, MiniCPM-o 4.5, Qwen2.5-Omni-7B, Gemma-4-E4B-it, Baichuan-Omni-1.5
Author countries
Australia, China, New Zealand