Spooftral: Can Voxtral Audio-Language Model Detect Speech Spoofing?

Authors: Avishai Weizman, Yehuda Ben-Shimol, Itshak Lapidot

Published: 2026-09-23 18:51:29+00:00

Comment: 8 pages, 3 figures, 5 tables. Accepted to the Spoken Language Technology (SLT) 2026

AI Summary

This study introduces Spooftral, an audio-language model (ALM) based on Voxtral, for speech spoofing detection. It proposes an instruction-guided approach using label-sequence likelihoods to differentiate bonafide from spoofed speech. The research finds that while raw LLM layers reduce spoof-discriminative cues, lightweight adaptation through DoRA significantly improves performance, achieving a 4.25% EER on the ASVspoof5 evaluation set.

Abstract

Self-supervised learning (SSL) countermeasures (CMs) have shown strong performance in recent years. However, they often show degraded performance while facing unseen spoofing attacks and mismatched conditions. This study examines the Voxtral audio-language model (ALM) framework for spoofing detection, as a step toward combining CM capabilities within the ALM framework. We analyze how Voxtral captures spoofing cues through audio-text processing and propose an instruction-guided approach that uses label-sequence likelihoods to evaluate bonafide and spoofed speech. Experiments on the ASVspoof databases show that without task-specific adaptation, the LLM layers emphasize semantic representations, reducing the separability of spoof-discriminative acoustic cues compared to the Whisper-based audio encoder. Consequently, spoofing-related information becomes less separable after language-model processing. We also applied lightweight adaptation using weight-decomposed low-rank adaptation (DoRA) to the Voxtral model and propose the Spooftral model, achieving an equal error rate (EER) of 4.25% on the ASVspoof5 evaluation set.


Key findings
Without task-specific adaptation, Voxtral's LLM layers reduce the separability of spoof-discriminative acoustic cues compared to the Whisper-based audio encoder. However, applying lightweight DoRA adaptation significantly improves performance, with Spooftral achieving an EER of 4.25% on the ASVspoof5 evaluation set, outperforming the audio-encoder-only variant and demonstrating that the LLM can contribute effectively to spoof detection when adapted.
Approach
The proposed Spooftral model adapts the Voxtral ALM for spoofing detection by formulating it as an instruction-guided generative label-likelihood classification task. It uses weight-decomposed low-rank adaptation (DoRA) to fine-tune the Voxtral model, specifically its attention layers, and employs length-normalized log-likelihood differences between "bonafide" and "spoof" label sequences for detection scores. A variant (Spooftral-Enc) excludes the LLM decoder for comparison.
Datasets
ASVspoof2019 LA, ASVspoof2021 LA, ASVspoof2021 DF, ASVspoof5
Model(s)
Voxtral-mini-3B (based on Ministral-3B), Whisper Large-v3 (as audio encoder), DoRA (weight-decomposed low-rank adaptation)
Author countries
Israel, France