Tracking the Trend in How Speech Synthesizers Deceive People
Authors: Milan Šalko, Anton Firc, Kamil Malinka, Vojtěch Staněk, Martin Perešini, Filip Pleško, Jakub Reš
Published: 2026-08-20 12:27:37+00:00
Comment: Accepted at the 6th Symposium on Security and Privacy in Speech Communication (SPSC 2026)
AI Summary
This study investigates human and automated deepfake audio detection capabilities against increasingly sophisticated speech synthesizers, specifically focusing on full and partial spoofing scenarios. It reveals a significant decline in human detection accuracy for newer, commercially available deepfake tools like ElevenLabs, even when listeners are warned, and highlights the increased deception of partial spoofs where a single sentence is altered. The research also benchmarks human performance against automated detectors, showing complementary failures and emphasizing the need for robust, layered defenses beyond human perception or single detectors.
Abstract
Advances in speech synthesis have made deepfake audio highly realistic. Earlier studies reported 70-80% human detection accuracy, but relied primarily on older synthesizers. We compare human detection for three selected voice synthesis tools released in 2019, 2022, and 2024 with 82 IT professionals, and benchmark humans against six pretrained detectors on the same material. For fully synthetic speech (full spoofs), the F1 score drops from about 90% for RTVC and YourTTS to 48% for ElevenLabs, although listeners were explicitly warned that deepfakes were present. For partial spoofing, where only one sentence of an utterance is altered, strict accuracy falls to 9%, and listeners classify the synthetic sentence as bona fide 77% of the time. Humans and detectors fail in complementary ways, and neither reliably localizes short manipulations. Additionally, listeners increasingly mislabel bona fide speech as fake, eroding trust in unmanipulated audio. These findings show that human perception alone is unreliable for the selected modern and partial-spoof conditions and motivate procedural verification, provenance, watermarking, and segment-level detection.