Large Audio Language Models for Spoofing-Aware Speaker Verification
Authors: Sofya Savelyeva, Mariia Perunova, Evgeny Kushnir, Artem Dvirniak, Dmitrii Korzh, Oleg Y. Rogov
Published: 2026-07-16 09:38:45+00:00
AI Summary
This research explores the application of Large Audio Language Models (LALMs) for Spoofing-Aware Speaker Verification (SASV), a task combining speaker verification with deepfake detection. The study evaluates LALMs under various conditions, including zero-shot prompting, supervised adaptation, and novel reasoning-oriented training and reinforcement learning optimizations. The findings demonstrate that while LALMs are not inherently suited for SASV in a zero-shot setting, task-specific adaptations effectively bridge this performance gap, yielding competitive results against conventional pipelines and highlighting their potential as an auditable foundation for unified SASV.
Abstract
Recent advances in text-to-speech and voice cloning make high-quality spoofing inexpensive and scalable, threatening voice authentication systems, especially automatic speaker verification (ASV). Existing defenses mainly address this threat through binary countermeasures (CMs) for deepfake detection or spoofing-aware speaker verification (SASV), where current systems are dominated by modular ASV-CM fusion and cascaded pipelines. Although large audio language models (LALMs) have shown promise on related audio tasks, including CM and ASV, their use for SASV remains unexplored, despite their capacity to produce natural-language rationales for auditing and robustness beyond discriminative predictions. This work systematically evaluates LALMs for SASV against conventional pipelines under zero-shot prompting, supervised adaptation, reasoning-oriented training, and reinforcement-learning-based optimization. Our results show that pretrained LALMs are near chance in the zero-shot setting, confirming that they are not natively suited to SASV, but that task-specific adaptation closes this gap. We further find that competitive SASV performance can be achieved through several distinct routes. These findings position LALMs as a promising and auditable foundation for unified SASV, while clarifying where conventional cascade systems still lead.