The Last Mile of Deepfake Speech Detection: An Industry-Academia Experience Report
Authors: Anton Firc, Kamil Malinka, Vojtěch Staněk, Miroslav Hlaváček, Marek Bartoň
Published: 2026-08-18 09:50:02+00:00
Comment: Accepted at the 6th Symposium on Security and Privacy in Speech Communication (SPSC 2026)
AI Summary
This industry-academia experience report highlights the significant challenges in deploying deepfake speech detectors despite impressive benchmark results. The authors, through a three-year project, identify barriers related to commercially usable datasets, realistic deployment benchmarks, and interpretable output scores for non-experts. The paper proposes research and coordination actions to bridge the gap between academic benchmarks and real-world deployment needs.
Abstract
Synthetic speech detection benchmarks now report sub-1% error rates on some in-domain evaluations, yet performance degrades under unseen attacks, channel mismatch, and distribution shift. Based on a three-year effort with Phonexia, a commercial speaker-recognition vendor, we report barriers encountered while building and deploying a detector. Many public benchmarks are not licensed for commercial model development. Real inputs are not four-second clean clips but long, codec-degraded, sometimes partially synthetic recordings. And when a calibrated system returns a log-likelihood ratio of 2.5, no one can tell the customer what it means for their decision. Rather than proposing a new model, we connect these barriers to concrete research and coordination proposals: shared standards for commercially usable datasets, realistic deployment benchmarks, and scores that non-experts can act on. These observations come from one project and should be tested in other settings.