What Survives the Codec Shift: Pooled No-Vocals Residuals for Speech Deepfake Detection
Authors: Jiajun Xu, Menglu Li, Xiao-Ping Zhang
Published: 2026-09-27 08:58:08+00:00
Comment: 5 pages, 2 figures, 3 tables. Prepared for submission to ICASSP 2027
AI Summary
This paper addresses the challenge of detecting speech deepfakes generated by neural codecs, which are more difficult to generalize against than vocoder-based fakes. The authors propose MN-P, a dual-view detector that combines token-level XLS-R features with an utterance-level pooled no-vocals residual representation, showing significant performance improvements, especially under unseen codec conditions. This approach demonstrates that complementary evidence from no-vocals residuals is crucial for robust cross-generation deepfake detection.
Abstract
The transition from vocoder-based to neural-codec speech synthesis makes generalization more difficult for speech deepfake detectors, particularly those relying on speech-oriented representations. It remains unclear which acoustic representations retain discriminative information when the generation mechanism changes. We therefore compare 12 acoustic representations using a shared low-capacity linear classifier to identify effective evidence under codec shift. The analysis shows that hierarchical XLS-R leads on the pooled test set, while pooled no-vocals residual statistics perform best on the unseen-codec condition, revealing complementary behavior across generation conditions. Building on this finding, we propose MN-P, a dual-view detector that integrates an utterance-level pooled no-vocals representation with token-level XLS-R features through adaptive gating. The proposed MN-P reduces EER by 54.2% overall and by 60.9% on the codec-unseen condition relative to the best-performing retrained state-of-the-art system, with consistent gains across different detector backends. These results indicate that pooled no-vocals residual statistics provide effective complementary evidence for cross-generation speech deepfake detection.