Spoofing detection under noisy conditions: a preliminary investigation and an initial database

Authors: Xiaohai Tian, Zhizheng Wu, Xiong Xiao, Eng Siong Chng, Haizhou Li

Published: 2016-02-09 12:00:56+00:00

Comment: Submitted to Odyssey: The Speaker and Language Recognition Workshop 2016

AI Summary

This paper presents a preliminary investigation into spoofing detection for automatic speaker verification (ASV) under additive noisy conditions, addressing a gap in previous research which primarily used clean data. The authors introduce a new noisy database, created by augmenting the ASVspoof 2015 database with five types of background noise at various signal-to-noise ratios (SNRs). Their experiments reveal that systems trained on clean data suffer significant performance degradation in noisy environments, with phase-based features showing greater robustness than magnitude-based ones.

Abstract

Spoofing detection for automatic speaker verification (ASV), which is to discriminate between live speech and attacks, has received increasing attentions recently. However, all the previous studies have been done on the clean data without significant additive noise. To simulate the real-life scenarios, we perform a preliminary investigation of spoofing detection under additive noisy conditions, and also describe an initial database for this task. The noisy database is based on the ASVspoof challenge 2015 database and generated by artificially adding background noises at different signal-to-noise ratios (SNRs). Five different additive noises are included. Our preliminary results show that using the model trained from clean data, the system performance degrades significantly in noisy conditions. Phase-based feature is more noise robust than magnitude-based features. And the systems perform significantly differ under different noise scenarios.


Key findings
The performance of spoofing detection systems degrades significantly in all noisy scenarios, with lower SNRs leading to worse results, and non-stationary noises generally impacting performance more severely. Even at 20 dB SNR, the best Equal Error Rate (EER) is about 4% lower than clean conditions. Phase-based features (e.g., Instantaneous Frequency, Baseband Phase Difference) demonstrated greater robustness to noise compared to magnitude-based features.
Approach
The authors construct a noisy database by artificially adding five different types of background noise (white, babble, Volvo, street, cafe) at 20dB, 10dB, and 0dB SNRs to the ASVspoof 2015 challenge database. They then evaluate a state-of-the-art spoofing detection system, comprising six feature extraction modules (LMS, RLMS, IF, BPD, GD, MGD) feeding into individual Multilayer Perceptron (MLP) classifiers, trained on clean data, with scores fused for a final decision.
Datasets
ASVspoof challenge 2015 database, NOISEX-92 database, QUT-NOISE database
Model(s)
Multilayer Perceptron (MLP) with one hidden layer of 2,048 sigmoid nodes
Author countries
Singapore, United Kingdom