Perceptible or Not? Diagnosing Passive Fingerprints for Speech Deepfake Attribution

Authors: Yupei Li, Qiyang Sun, Emmanouil Benetos, Berrak Sisman, Björn Schuller

Published: 2026-09-01 05:54:02+00:00

AI Summary

This paper introduces the Perceptible-Imperceptible Passive-fingerprint Diagnostic Protocol (PIPDP) to analyze the reliability of perceptible and imperceptible passive fingerprints for speech deepfake attribution. Experiments with ten speech generators and three attribution detectors demonstrate that imperceptible fingerprints offer more persistent and reliable attribution cues compared to perceptible ones, which are less stable and easily manipulated.

Abstract

Passive fingerprints (intrinsic traces naturally left by generators) have been shown to enable attribution in speech deepfake detection, yet their persistence, reproducibility, and content-independence remain unverified. Moreover, no prior work distinguishes perceptible from imperceptible fingerprints, although the two have very different implications for attribution reliability. Perceptible fingerprints, such as emotional expression, are shaped by perceptual quality objectives and may change across model updates, whereas imperceptible fingerprints are not explicitly optimised by current training objectives and are rarely considered in existing dataset design or training strategies, as they have limited influence on downstream applications. We therefore propose a Perceptible-Imperceptible Passive-fingerprint Diagnostic Protocol (PIPDP) to define and separately analyze these two fingerprint types. PIPDP comprises three complementary analyses: multi-evidence fingerprint verification through residual-energy, reproducibility, and saliency analyses, perceptually transparent perturbations preserving audio quality, and prompt-driven emotion change that modifies perceptible fingerprints without model retraining. Experiments across ten speech generators and three attribution detectors show that imperceptible fingerprints provide persistent attribution cues. Perceptually transparent perturbations reduce attribution accuracy by up to 48.2\\% on HiggsAudioV3, whereas emotion-driven changes leave attribution largely unchanged, with only about a 1.0\\% accuracy variation across emotions on CosyVoice2 using w2v-bert-MLP. These results suggest that imperceptible fingerprints are more reliable for trustworthy attribution.


Key findings
Perceptually transparent perturbations significantly reduce attribution accuracy (up to 48.2% on HiggsAudioV3), indicating the reliance of attribution on imperceptible fingerprints. Conversely, emotion-driven changes (modifying perceptible fingerprints) result in minimal accuracy variation (around 1.0% on CosyVoice2), suggesting their unreliability for consistent attribution. These findings establish imperceptible fingerprints as more reliable for trustworthy speech deepfake attribution.
Approach
The authors propose PIPDP, a three-probe diagnostic protocol. Probe 1 verifies fingerprint existence and reproducibility through residual-energy, reproducibility, and saliency analyses. Probe 2 assesses whether reliable fingerprints are imperceptible using perceptually transparent audio perturbations. Probe 3 evaluates the stability of perceptible fingerprints by modifying emotion via prompt-driven generation.
Datasets
VCTK corpus, ESD dataset
Model(s)
WavLM-AASIST, RawNet2, w2v-bert-MLP
Author countries
UK, USA, Germany