Tracing and Relearning Detection Evidence in Text-to-Speech Systems

Authors: Eunji Shin, Kyudan Jung, Jihwan Kim, Minwoo Lee, Jaegul Choo

Published: 2026-09-25 08:29:14+00:00

Comment: Submitted to ICASSP 2027

AI Summary

This research investigates which stage of a text-to-speech (TTS) system contributes most to audio deepfake detection evidence, using a controlled resynthesis and detector adaptation approach. They find that acoustic generation introduces more detectable evidence than vocoder reconstruction. Although fine-tuning the acoustic model can reduce the detectability of synthetic speech by fixed detectors, adapting the detectors to these updated models can successfully re-establish detectability.

Abstract

Recent audio deepfake detectors separate bona fide speech from synthetic speech, yet it remains unclear which stage of a text-to-speech system supplies the detection evidence. We address this with controlled resynthesis and detector adaptation in an F5-TTS-BigVGAN pipeline. Since vocoder reconstruction of a real mel can itself be separable from the source utterance, we fix the vocoder and trace the larger change in detector separation to acoustic generation. Adversarially fine-tuning the acoustic model, with no detector in its objective, raises EER against fixed detectors at comparable quality. However, adapting a detector only on the tuned model's VCTK outputs lowers its LibriSpeech EER from 19.42% to 7.46% and improves detection of unseen base F5-TTS outputs. These results suggest that acoustic-model updates can reduce the detection evidence available to fixed detectors, while detector adaptation keeps the updated outputs detectable in this pipeline.


Key findings
Acoustic generation, rather than vocoder reconstruction, provides the larger portion of deepfake detection evidence in the F5-TTS–BigVGAN pipeline. Fine-tuning the acoustic model can increase the Equal Error Rate (EER) against fixed detectors, making synthetic speech harder to detect without degrading perceived quality. However, adapting detectors to these newly tuned outputs effectively lowers EERs, demonstrating that detection evidence is still present and can be relearned.
Approach
The authors analyze a F5-TTS–BigVGAN pipeline, fixing the vocoder to isolate the contribution of acoustic generation to detection evidence. They fine-tune the acoustic model adversarially without a detector in its objective, then adapt existing deepfake detectors to the outputs of this fine-tuned model. This process helps trace where detection evidence originates and how it changes with model updates.
Datasets
LibriSpeech, VCTK, VoxPopuli, ASVspoof 2019, ASVspoof 2021 DF/LA, In-the-Wild
Model(s)
F5-TTS (Flow-matching Diffusion Transformer acoustic model), BigVGAN, Vocos, XLS-R SLS, XLSR-Mamba, AASIST-L
Author countries
South Korea