Tracing and Relearning Detection Evidence in Text-to-Speech Systems
Authors: Eunji Shin, Kyudan Jung, Jihwan Kim, Minwoo Lee, Jaegul Choo
Published: 2026-09-25 08:29:14+00:00
Comment: Submitted to ICASSP 2027
AI Summary
This research investigates which stage of a text-to-speech (TTS) system contributes most to audio deepfake detection evidence, using a controlled resynthesis and detector adaptation approach. They find that acoustic generation introduces more detectable evidence than vocoder reconstruction. Although fine-tuning the acoustic model can reduce the detectability of synthetic speech by fixed detectors, adapting the detectors to these updated models can successfully re-establish detectability.
Abstract
Recent audio deepfake detectors separate bona fide speech from synthetic speech, yet it remains unclear which stage of a text-to-speech system supplies the detection evidence. We address this with controlled resynthesis and detector adaptation in an F5-TTS-BigVGAN pipeline. Since vocoder reconstruction of a real mel can itself be separable from the source utterance, we fix the vocoder and trace the larger change in detector separation to acoustic generation. Adversarially fine-tuning the acoustic model, with no detector in its objective, raises EER against fixed detectors at comparable quality. However, adapting a detector only on the tuned model's VCTK outputs lowers its LibriSpeech EER from 19.42% to 7.46% and improves detection of unseen base F5-TTS outputs. These results suggest that acoustic-model updates can reduce the detection evidence available to fixed detectors, while detector adaptation keeps the updated outputs detectable in this pipeline.