VietPrism: A large-scale Vietnamese speech and deepfake corpus with diverse dialects and code-switching

Authors: Minh Hoang, Thai Le

Published: 2026-09-24 15:48:23+00:00

Comment: Preprint for ICASSP 2027 submission

AI Summary

The authors introduce VietPrism, a large-scale, multi-domain Vietnamese speech and deepfake corpus, featuring 993.4 hours of bona fide speech and over 3.1K hours of spoof speech. This corpus uniquely combines transcripts, consistent speaker identities, five dialect groups, and Vietnamese-English code-switching, enabling controlled evaluation of audio deepfake detection. Initial zero-shot evaluations on multilingual detectors reveal significant brittleness, with performance varying greatly across generators, speaker similarity levels, and dialects.

Abstract

Vietnamese speech research is constrained by resources that isolate automatic speech recognition from speaker, dialect, code-switching, and deepfake analysis. We introduce VietPrism, an open, multi-domain corpus that brings these dimensions together at scale: 993.4 hours and 403,941 bona fide utterances from 1,262 verified speakers across 8,388 real-world videos. To our knowledge, it is the first large-scale Vietnamese corpus to jointly provide transcripts, consistent speaker identities, five dialect groups, and naturally occurring Vietnamese--English code-switching, which constitutes nearly half of the corpus by duration. We further create over 3.1K hours of spoof speech with four open-source and commercial synthesis systems. Every spoof is conditioned on a verified speaker reference and paired with a transcript- and speaker-matched bona fide utterance, enabling unique controlled evaluation with reduced lexical and identity confounds. Zero-shot evaluation of five pretrained multilingual detectors reveals striking brittleness: EER greatly varies across detector--generator pairings, while recent multilingual detector DFA-1B degrades from 16.3% to 33.6% as speaker similarity increases. Dialect-stratified results expose further model-dependent disparities. By unifying natural linguistic diversity with controlled spoof generation, VietPrism provides a challenging foundation for Vietnamese speech modeling and trustworthy audio-deepfake detection.


Key findings
Zero-shot evaluation of multilingual deepfake detectors on VietPrism shows significant brittleness, with EER widely varying based on the specific detector-generator pairing. Detector performance degrades significantly with increased speaker similarity between spoofed and real speech, and also exhibits model-dependent disparities across different Vietnamese dialects. Commercial synthesizers, like MiniMax, demonstrate higher resistance to detection under code-switching conditions compared to open-source models.
Approach
The authors developed a comprehensive data generation pipeline that includes video discovery and curation, speech processing and alignment, transcript quality control, and speaker-conditioned spoof generation. They collected real-world Vietnamese videos, meticulously processed the audio, generated spoof audio using various synthesis systems while maintaining speaker and transcript matching, and then evaluated several pre-trained multilingual deepfake detectors on this new corpus in a zero-shot setting.
Datasets
VietPrism (newly introduced), VIVOS, FPT Open Speech, Common Voice, Bud500, VietSpeech, GigaSpeech 2, Vietnam-Celeb, VoxVietnam, CanVEC, ViMedCSS, MLAAD, VSASV, JMAD, SF-MD, SEA-Spoof (for comparison)
Model(s)
DF-Arena (DFA) 1B, DF-Arena (DFA) 500M, AntiDeepfake (ADF) (W2V2-L, XLS-R-2B, MMS-300M)
Author countries
UNKNOWN, USA