Is Broader Better? A Controlled Study of Multilingual Coverage and Pretraining Objective in Frozen SSL Encoders for Speech Deepfake Detection

Authors: Benjamin Hurt, Oscar O'Donnell

Published: 2026-09-24 07:14:12+00:00

Comment: 5 pages, 1 figure, 2 tables. Submitted to ICASSP 2027

AI Summary

This research investigates the impact of multilingual coverage and pretraining objectives in frozen self-supervised (SSL) speech encoders for audio deepfake detection. The study reveals that increasing multilingual coverage beyond approximately 100 languages does not monotonically improve out-of-domain generalization, and masked-prediction pretraining generalizes significantly better than contrastive pretraining on identical data. The findings suggest that broader and larger models are not always better for this task.

Abstract

Frozen self-supervised (SSL) speech encoders are strong, low-cost front ends for audio deepfake detection, and recent comparisons agree that large, multilingual, discriminative encoders generalize best out of domain. These comparisons fail to control for encoder capacity, pretraining objective, and multilingual coverage together, identifying which encoder wins without isolating why. We present a controlled decomposition with a fixed pipeline and trainable capacity. We vary multilingual coverage on four wav2vec2-family encoders, matched to ~315M parameters. We isolate the pretraining objective on two encoders matched on identical data. Coverage does not help monotonically, as out-of-domain error drops sharply at the ~100-language scale (XLS-R) but does not improve further at the 1406-language extreme (MMS). We find that a mid-coverage encoder is strongest on farther out-of-domain sets, matching or surpassing a 577M-parameter model at 315M. Its lead on these far sets, statistically significant under paired bootstrap, and on the official ASVspoof 5 cost metric holds under two backends. Separately, masked-prediction pretraining generalizes better than contrastive on identical data (In-the-Wild EER 26.5% vs. 46.8%). Within this fixed frozen-encoder recipe, we find that broader and larger models are not reliably better.


Key findings
Multilingual coverage shows a non-monotonic effect on out-of-domain generalization, with performance peaking around 100 languages (XLS-R) and not improving further with 1406 languages (MMS). Masked-prediction pretraining (HuBERT) significantly outperforms contrastive pretraining (wav2vec2-LV60) on identical data. A 315M-parameter encoder (XLS-R) can match or surpass a 577M-parameter model (XEUS) on farther out-of-domain sets, indicating that broader and larger models are not reliably better.
Approach
The authors perform a controlled decomposition by fixing a detection pipeline and trainable capacity, varying multilingual coverage across four wav2vec2-family encoders matched at ~315M parameters, and isolating the pretraining objective on two encoders with identical training data. They evaluate performance on various out-of-domain benchmarks using pooled Equal Error Rate (EER) and ASVspoof 5 min-DCF with two different backends.
Datasets
ASVspoof 2019 LA, ASVspoof 2021 (LA21, DF21), In-the-Wild (ITW), ASVspoof 5 (ASV5)
Model(s)
wav2vec2-LV60, XLSR-53, XLS-R, MMS-300M, HuBERT-LARGE, WavLM-LARGE, XEUS (reference), MHFA (backend), AASIST (backend)
Author countries
United States