Neural Audio Codec for Robust Audio Deepfake Detection

Authors: Jungwoo Kim, Joonyong Park, Junyoung Koh, Jong-Seok Lee

Published: 2026-09-30 12:53:53+00:00

Comment: 5 pages, 7 figures

AI Summary

This research investigates the impact of audio coding on deepfake detection, finding that low bitrates significantly degrade performance. The authors propose a forensic-preserving neural audio codec (FP-NAC) that fine-tunes a pre-trained codec with a detector-guided objective to improve deepfake detection robustness while maintaining reconstruction quality.

Abstract

Audio deepfake detectors are typically evaluated on uncompressed audio, although real-world audio often undergoes low-bitrate coding. In this work, we investigate how audio coding affects deepfake detection across codecs, bitrates, and detectors, finding higher errors at lower rates. A mixed-pair protocol isolates codec-induced changes in bona fide and spoof audio, revealing asymmetric, codec-dependent failures: low-rate DAC and EnCodec mainly degrade bona fide detection, whereas X-Codec shows a stronger spoof-side limitation. Motivated by these, we propose a forensic-preserving neural audio codec (FP-NAC), which fine-tunes a pretrained codec using a detector-guided objective while preserving its native hard quantization path and bitrate. On ASVspoof 2019 LA, FP-NAC reduces EER by up to 49.8~pp compared with the original DAC at 0.5~kbps while maintaining comparable reconstruction quality. Although supervised by only one detector, FP-NAC improves performance across multiple detectors, highlighting forensic transparency as a codec design objective alongside perceptual quality. Our codes are available at https://github.com/kjungwoo03/FP-NAC.


Key findings
FP-NAC significantly reduces Equal Error Rate (EER) by up to 49.8 percentage points at 0.5 kbps compared to the original DAC, demonstrating improved robustness against coding artifacts. These improvements transfer consistently to held-out detectors (AASIST-L, RawNet2) even though only AASIST was used for supervision, and the codec maintains comparable reconstruction quality.
Approach
The authors analyze deepfake detection failures under various audio codecs and bitrates using a mixed-pair protocol. They then propose FP-NAC, which adapts a pre-trained neural audio codec (DAC) by integrating K-conditioned residual adapters and optimizing with a forensic-preserving objective, guided by a frozen deepfake detector, to maintain forensic cues.
Datasets
ASVspoof 2019 LA
Model(s)
DAC (base codec), AASIST, AASIST-L, RawNet2
Author countries
Republic of Korea, Japan, USA