VietPrism: A large-scale Vietnamese speech and deepfake corpus with diverse dialects and code-switching
Authors: Minh Hoang, Thai Le
Published: 2026-09-24 15:48:23+00:00
Comment: Preprint for ICASSP 2027 submission
AI Summary
The authors introduce VietPrism, a large-scale, multi-domain Vietnamese speech and deepfake corpus, featuring 993.4 hours of bona fide speech and over 3.1K hours of spoof speech. This corpus uniquely combines transcripts, consistent speaker identities, five dialect groups, and Vietnamese-English code-switching, enabling controlled evaluation of audio deepfake detection. Initial zero-shot evaluations on multilingual detectors reveal significant brittleness, with performance varying greatly across generators, speaker similarity levels, and dialects.
Abstract
Vietnamese speech research is constrained by resources that isolate automatic speech recognition from speaker, dialect, code-switching, and deepfake analysis. We introduce VietPrism, an open, multi-domain corpus that brings these dimensions together at scale: 993.4 hours and 403,941 bona fide utterances from 1,262 verified speakers across 8,388 real-world videos. To our knowledge, it is the first large-scale Vietnamese corpus to jointly provide transcripts, consistent speaker identities, five dialect groups, and naturally occurring Vietnamese--English code-switching, which constitutes nearly half of the corpus by duration. We further create over 3.1K hours of spoof speech with four open-source and commercial synthesis systems. Every spoof is conditioned on a verified speaker reference and paired with a transcript- and speaker-matched bona fide utterance, enabling unique controlled evaluation with reduced lexical and identity confounds. Zero-shot evaluation of five pretrained multilingual detectors reveals striking brittleness: EER greatly varies across detector--generator pairings, while recent multilingual detector DFA-1B degrades from 16.3% to 33.6% as speaker similarity increases. Dialect-stratified results expose further model-dependent disparities. By unifying natural linguistic diversity with controlled spoof generation, VietPrism provides a challenging foundation for Vietnamese speech modeling and trustworthy audio-deepfake detection.