Language Orthogonalization for Zero-Shot Cross-Lingual Audio Deepfake Detection
Authors: Minu Kim, Ji Sub Um, Hoirin Kim
Published: 2026-09-15 00:23:42+00:00
Comment: Submitted to ICASSP 2027
AI Summary
This research addresses the challenge of zero-shot cross-lingual audio deepfake detection, where detectors need to generalize to languages not seen during training. The authors propose language orthogonalization, a target-free ridge map method that removes language-dependent variations from self-supervised speech model (S3M) representations, thereby isolating spoof cues. This approach consistently reduces Equal Error Rate (EER) across various unseen languages and S3M backbones, with greater gains for more linguistically distant transfers.
Abstract
Audio deepfake detectors need to transfer to languages absent from training, as multilingual speech synthesis outpaces labeled anti-spoofing resources. While detectors increasingly rely on self-supervised speech models (S3Ms), these backbones encode language-dependent structure that confounds spoof cues. We address this confound through language orthogonalization, a target-free ridge map that removes S3M variation projected onto continuous language-identification (LID) embeddings. Across six languages, six S3M backbones, and all Leave-N-Out settings, it consistently reduces EER across unseen languages. Cross-lingual EER correlates with LID-space distance, where orthogonalization yields larger gains for more distant transfers.