The Last Mile of Deepfake Speech Detection: An Industry-Academia Experience Report

Authors: Anton Firc, Kamil Malinka, Vojtěch Staněk, Miroslav Hlaváček, Marek Bartoň

Published: 2026-08-18 09:50:02+00:00

Comment: Accepted at the 6th Symposium on Security and Privacy in Speech Communication (SPSC 2026)

AI Summary

This industry-academia experience report highlights the significant challenges in deploying deepfake speech detectors despite impressive benchmark results. The authors, through a three-year project, identify barriers related to commercially usable datasets, realistic deployment benchmarks, and interpretable output scores for non-experts. The paper proposes research and coordination actions to bridge the gap between academic benchmarks and real-world deployment needs.

Abstract

Synthetic speech detection benchmarks now report sub-1% error rates on some in-domain evaluations, yet performance degrades under unseen attacks, channel mismatch, and distribution shift. Based on a three-year effort with Phonexia, a commercial speaker-recognition vendor, we report barriers encountered while building and deploying a detector. Many public benchmarks are not licensed for commercial model development. Real inputs are not four-second clean clips but long, codec-degraded, sometimes partially synthetic recordings. And when a calibrated system returns a log-likelihood ratio of 2.5, no one can tell the customer what it means for their decision. Rather than proposing a new model, we connect these barriers to concrete research and coordination proposals: shared standards for commercially usable datasets, realistic deployment benchmarks, and scores that non-experts can act on. These observations come from one project and should be tested in other settings.


Key findings
Key findings indicate that current benchmarks do not reflect real-world deployment challenges, particularly regarding generalization to unseen attacks, channel mismatch, and data drift. There's a critical need for commercially usable datasets with clear licensing, evaluation metrics that reflect operational realities, and methods to make detector outputs intelligible and actionable for non-expert users. Addressing these issues requires both technical research and community-wide coordination on standards and shared resources.
Approach
This paper is an experience report rather than proposing a new model. It analyzes barriers encountered during the deployment of a deepfake speech detector and outlines open problems and proposed actions for the community. The authors' internal development, however, involved pretrained self-supervised (SSL) front-ends with attentive pooling.
Datasets
Not applicable (experience report). The paper discusses limitations of public benchmarks like ASVspoof 2019, ASVspoof 2021, ASVspoof 5, and mentions commercial synthesizers like ElevenLabs.
Model(s)
Not applicable (experience report focused on deployment challenges). The authors mention working with open-source architectures and their own university-contributed models, settling on pretrained SSL front-ends with attentive pooling.
Author countries
Czech Republic