What Survives the Codec Shift: Pooled No-Vocals Residuals for Speech Deepfake Detection

Authors: Jiajun Xu, Menglu Li, Xiao-Ping Zhang

Published: 2026-09-27 08:58:08+00:00

Comment: 5 pages, 2 figures, 3 tables. Prepared for submission to ICASSP 2027

AI Summary

This paper addresses the challenge of detecting speech deepfakes generated by neural codecs, which are more difficult to generalize against than vocoder-based fakes. The authors propose MN-P, a dual-view detector that combines token-level XLS-R features with an utterance-level pooled no-vocals residual representation, showing significant performance improvements, especially under unseen codec conditions. This approach demonstrates that complementary evidence from no-vocals residuals is crucial for robust cross-generation deepfake detection.

Abstract

The transition from vocoder-based to neural-codec speech synthesis makes generalization more difficult for speech deepfake detectors, particularly those relying on speech-oriented representations. It remains unclear which acoustic representations retain discriminative information when the generation mechanism changes. We therefore compare 12 acoustic representations using a shared low-capacity linear classifier to identify effective evidence under codec shift. The analysis shows that hierarchical XLS-R leads on the pooled test set, while pooled no-vocals residual statistics perform best on the unseen-codec condition, revealing complementary behavior across generation conditions. Building on this finding, we propose MN-P, a dual-view detector that integrates an utterance-level pooled no-vocals representation with token-level XLS-R features through adaptive gating. The proposed MN-P reduces EER by 54.2% overall and by 60.9% on the codec-unseen condition relative to the best-performing retrained state-of-the-art system, with consistent gains across different detector backends. These results indicate that pooled no-vocals residual statistics provide effective complementary evidence for cross-generation speech deepfake detection.


Key findings
Pooled no-vocals residual statistics provide effective complementary evidence for cross-generation speech deepfake detection, especially in codec-unseen scenarios. The proposed MN-P detector significantly reduces EER by 54.2% overall and 60.9% on codec-unseen conditions compared to the best retrained state-of-the-art systems. Aggregating no-vocals residuals at an utterance level is more effective than using them as a frame-level sequence.
Approach
The approach involves a two-stage process. First, an analysis of 12 acoustic representations reveals that hierarchical XLS-R excels generally, while pooled no-vocals residual statistics perform best under unseen codec conditions. Second, the MN-P detector is proposed, which integrates these two complementary representations: token-level XLS-R features and an utterance-level pooled no-vocals residual, using adaptive gating for fusion.
Datasets
ASVspoof2019-LA, Codecfake
Model(s)
XLS-R-300M, AASIST, Nes2Net, HTDemucs (for no-vocals separation)
Author countries
China