WST-Graph: Topology-Preserving Wavelet Scattering Front-End for Speech Deepfake Detection

Authors: Kwok-Ho Ng, Tingting Song, Bingwen Feng, Zhihua Xia

Published: 2026-09-24 10:53:14+00:00

AI Summary

This paper introduces WST-Graph, a novel front-end for speech deepfake detection that leverages the Wavelet Scattering Transform (WST) to preserve the inherent carrier-modulation topology of acoustic features. By reconstructing WST paths as a sparse grid and integrating it with an AASIST graph backend, WST-Graph aims to improve detection performance, especially on out-of-domain data, with fewer parameters. The approach emphasizes retaining explicit acoustic axes through modulation-level normalization and length-aware adaptive local attention pooling.

Abstract

The acoustic front-end determines which forensic cues a speech deepfake detector can exploit. The wavelet scattering transform (WST) provides stable multiscale coefficients with explicit coordinates, yet direct flattening obscures the parent relation between paths. We introduce WST-Graph, reconstructing these paths as a sparse modulation-carrier grid for an AASIST graph backend. Modulation-level normalization and length-aware adaptive local attention pooling produce fixed relative-time representations while retaining the acoustic axes before learned adaptation. This yields a waveform-to-graph interface with a fixed, parameter-free WST. Our configurations remain competitive with AASIST while using approximately 60% fewer trainable parameters and show clear gains on selected out-of-domain benchmarks. These results underscore the value of preserving parent-child relations within the carrier-modulation topology when constructing a compact, physically grounded interface for graph-based speech deepfake detection. Code will be released at https://github.com/saki-ciallo/wst-graph.


Key findings
WST-Graph configurations achieved competitive performance with AASIST while using approximately 60% fewer trainable parameters. It demonstrated clear gains on selected out-of-domain benchmarks, particularly on DFADD, underscoring the importance of preserving carrier–modulation topology. Optimal performance was observed with modulation-level normalization and attentive pooling for graph node construction.
Approach
WST-Graph reconstructs WST paths into a sparse modulation-carrier grid, preserving parent-child relations. It uses modulation-level normalization and length-aware adaptive local attention pooling (ALAP) to generate fixed relative-time representations. This processed output then interfaces with an AASIST-style graph backend for deepfake classification.
Datasets
ASVspoof 2019 LA, Speech DF Arena (including 14 distinct sets, 13 out-of-domain)
Model(s)
WST (Wavelet Scattering Transform), AASIST (Audio Anti-Spoofing using Integrated Spectro-Temporal Graph Attention Networks) graph backend
Author countries
China