This survey provides a comprehensive overview of AI voice generation and detection methods, addressing the rapid advancement of realistic AI-generated voices and the urgent need for robust safeguards against harmful uses like impersonation and fraud. It covers technical foundations, state-of-the-art advances, key open challenges, benchmark resources, and future directions in this evolving field.
The ability of artificial intelligence (AI) models to generate highly realistic human voices has advanced rapidly. These technologies power accessibility tools, virtual assistants and creative applications, but they also enable harmful uses, including impersonation, fraud and disinformation. Recent incidents of voice cloning scams targeting businesses and political leaders underscore the urgent need for robust safeguards. Unlike image and video deepfakes, the detection of synthetic voices poses unique challenges due to the complexity of phonetics, prosody and auditory perception. This survey offers a comprehensive overview of AI voice generation and detection methods, encompassing both the technical foundations and the latest state-of-the-art advances. This study also identifies key open challenges, benchmark resources and future directions to make this survey useful for future researchers.
Key findings
The increasing realism of AI-generated voices presents significant challenges for detection, necessitating a deeper understanding of both generation mechanisms and human speech physiology. While signal-based detection is fast and model-agnostic, phonetic and linguistic analyses offer more semantically grounded and robust cues against evolving synthetic speech. Hybrid systems combining both approaches and addressing domain shift are crucial for future robust detection.
Approach
This paper is a survey, not a research paper presenting a novel approach. It categorizes existing AI voice generation into Text-to-Speech (TTS), Voice Conversion (VC), and Voice Cloning, detailing their evolution and technical underpinnings. For detection, it discusses two main tracks: signal-based methods focusing on acoustic anomalies and phonetics-based analysis identifying linguistic inconsistencies.
Datasets
ASVspoof (2015, 2017, 2019, 2021, 5 2025), ReMASC 2019, FoR 2019, WaveFake 2021, FakeAVCeleb 2021, ADD 2022, LibriSeVoc 2023.
Model(s)
Not applicable, as this is a survey paper reviewing various models. However, the survey mentions various models/architectures in the context of generation and detection, including HMMs, GMMs, WaveNet, Tacotron, VALL-E, FastSpeech, Glow-TTS, VITS, PortaSpeech, Meta-StyleSpeech, NaturalSpeech, YourTTS, StyleTTS, BASE TTS, MaskGCT, CLaM-TTS, CM-TTS, XTTS, SupertonicTTS, VITA-Audio, DiFlow-TTS, AutoVC, CycleGAN-VC, StarGANv2-VC, VQMIVC, AVQVC, DRVC, NVC-Net, DDDM-VC, ProDiff, DiffGAN-TTS, UnitSpeech, ElevenLabs, CQCC-GMM, LFCC-GMM, LFCC-LCNN, RawNet2, AASIST, RawNet3, WavLM, XLS-R, RawNet2-vocoder.
Author countries
USA