Tracking the Trend in How Speech Synthesizers Deceive People

Authors: Milan Šalko, Anton Firc, Kamil Malinka, Vojtěch Staněk, Martin Perešini, Filip Pleško, Jakub Reš

Published: 2026-08-20 12:27:37+00:00

Comment: Accepted at the 6th Symposium on Security and Privacy in Speech Communication (SPSC 2026)

AI Summary

This study investigates human and automated deepfake audio detection capabilities against increasingly sophisticated speech synthesizers, specifically focusing on full and partial spoofing scenarios. It reveals a significant decline in human detection accuracy for newer, commercially available deepfake tools like ElevenLabs, even when listeners are warned, and highlights the increased deception of partial spoofs where a single sentence is altered. The research also benchmarks human performance against automated detectors, showing complementary failures and emphasizing the need for robust, layered defenses beyond human perception or single detectors.

Abstract

Advances in speech synthesis have made deepfake audio highly realistic. Earlier studies reported 70-80% human detection accuracy, but relied primarily on older synthesizers. We compare human detection for three selected voice synthesis tools released in 2019, 2022, and 2024 with 82 IT professionals, and benchmark humans against six pretrained detectors on the same material. For fully synthetic speech (full spoofs), the F1 score drops from about 90% for RTVC and YourTTS to 48% for ElevenLabs, although listeners were explicitly warned that deepfakes were present. For partial spoofing, where only one sentence of an utterance is altered, strict accuracy falls to 9%, and listeners classify the synthetic sentence as bona fide 77% of the time. Humans and detectors fail in complementary ways, and neither reliably localizes short manipulations. Additionally, listeners increasingly mislabel bona fide speech as fake, eroding trust in unmanipulated audio. These findings show that human perception alone is unreliable for the selected modern and partial-spoof conditions and motivate procedural verification, provenance, watermarking, and segment-level detection.


Key findings
Human detection of fully synthetic speech from ElevenLabs dropped significantly to an F1 score of 48%, compared to 90% for older tools, even with explicit warnings. Partial spoofing proved even more deceptive, with strict accuracy falling to 9%, and listeners misclassifying synthetic sentences as bona fide 77% of the time. Neither humans nor the evaluated automated detectors reliably localize short manipulations, suggesting complementary failures and the inadequacy of single-layer defenses against modern deepfake audio.
Approach
The authors conducted a questionnaire-based survey with IT professionals to assess their ability to detect deepfake audio generated by three synthesizers (RTVC, YourTTS, ElevenLabs) from different release years. They evaluated both fully synthetic speech and partial spoofs (one synthetic sentence in an otherwise bona fide utterance). Human performance was then benchmarked against six pretrained automated deepfake speech detectors on the same material.
Datasets
Custom-generated audio dataset using RTVC (2019), YourTTS (2022), and ElevenLabs (2024) based on well-known celebrity voices from YouTube interviews. Automated detectors were trained on ASVspoof 2019 LA and ASVspoof 5.
Model(s)
For automated detection: XLS-R 300M (frontend) + AASIST (backend), XLS-R 300M (frontend) + MHFA (backend), WavLM Base+ (frontend) + AASIST (backend), WavLM Base+ (frontend) + MHFA (backend).
Author countries
Czech Republic