On Evaluating Quantum Kernel Robustness for Low-Resource Cross-Corpus Audio Deepfake Detection

Authors: Lisan Al Amin, Lei Zhang, Vandana P. Janeja

Published: 2026-09-30 19:51:25+00:00

AI Summary

This study explores quantum kernel methods and lightweight neural models for low-resource, cross-corpus audio deepfake detection using a strict budget of 200 training samples and frozen wav2vec 2.0 embeddings. The research demonstrates that a Quantum Support Vector Machine (QSVM) can maintain meaningful discrimination under severe domain shifts where a Multilayer Perceptron (MLP) degrades to near-random performance, though this advantage is not consistent across all transfer directions. The findings characterize the inductive bias of quantum kernel methods under distribution shift rather than claiming quantum advantage, as the four-qubit kernel used is classically simulable.

Abstract

Synthetic speech detection is critical for audio security, but performance can degrade when labeled data are scarce and evaluation conditions differ from training. This study examines quantum kernel methods and lightweight neural models for cross-corpus audio deepfake detection under limited training data. We compare a Quantum Support Vector Machine (QSVM), a classical support vector machine (SVM), and a multilayer perceptron (MLP), all trained on frozen wav2vec 2.0 embeddings using a strict budget of 200 training samples. To match the qubit budget of near-term quantum hardware, embeddings are reduced to four dimensions using principal component analysis, and all models use the same reduced features. Experiments on ASVspoof 2019, ASVspoof 5, the ADD 2023 Challenge, and the In-the-Wild dataset show that under severe domain shift from ASVspoof 2019 to ADD 2023, the MLP degrades to near-random performance, with an area under the curve of approximately 50% and an equal error rate of 50.0%. In contrast, the QSVM maintains meaningful discrimination, achieving an area under the curve of 76.0% and an equal error rate of 27.0%. This advantage is not consistent across transfer directions. When trained on ADD 2023, the QSVM falls below chance on two of three transfers, while the MLP performs better. These results suggest that quantum kernel methods can be competitive under severe cross-corpus shifts and strict low-resource constraints, but do not provide a consistent advantage under near-domain transfer. We interpret these findings as an empirical characterization of quantum kernel inductive bias under distribution shift, rather than evidence of quantum advantage, since the four-qubit kernel can be simulated exactly on classical hardware.


Key findings
Under severe domain shift (ASVspoof 2019 to ADD 2023), the MLP degrades to near-random performance (AUC ~50%), while the QSVM maintains significant discrimination (AUC 76.0%, EER 27.0%). This advantage is not universal, as the QSVM can perform below chance on some transfers when trained on ADD 2023, while the MLP performs better. The study suggests that quantum kernel methods can be competitive for low-resource, cross-corpus deepfake detection, but their benefits are shift-dependent and stem from their inductive bias, not quantum advantage.
Approach
The authors utilize frozen wav2vec 2.0 embeddings, which are then reduced to four dimensions using Principal Component Analysis (PCA) to match the qubit budget of near-term quantum hardware. These reduced features are fed into three models: a Quantum Support Vector Machine (QSVM), a classical Support Vector Machine (SVM) with an RBF kernel, and a Multilayer Perceptron (MLP). All models are trained with a strict budget of 200 labeled training samples.
Datasets
ASVspoof 2019, ASVspoof 5, ADD 2023 Challenge, In-the-Wild dataset
Model(s)
Quantum Support Vector Machine (QSVM), Support Vector Machine (SVM) with RBF kernel, Multilayer Perceptron (MLP)
Author countries
USA