SEAR: Spoofing Evidence-Grounded Audio Reasoning Benchmark for Audio Language Models

Authors: Rong Wan, Suliu Qin, Jiaxi Li, Wei Xie, Wenwu Wang, Xiaolong Han, Lu Yin, Xilu Wang

Published: 2026-09-30 14:39:47+00:00

AI Summary

This paper introduces SEAR, a novel four-task Audio Question Answering (AQA) benchmark for evaluating Audio Language Model (ALM)-based deepfake detection by focusing on acoustic evidence identification, quantification, deepfake detection, and forensic rationale generation. It also proposes BAEA, a bona-fide-based acoustic evidence agent that augments a frozen ALM with controlled acoustic tools under fixed or adaptive evidence-acquisition policies. Experiments show a clear disparity between plausible rationales and verifiable acoustic evidence reasoning, with BAEA-FIXED significantly improving verdicts and rationales.

Abstract

Audio language models (ALMs) are increasingly used for audio deepfake detection (ADD), yet existing benchmarks assess their verdicts or rationale plausibility without verifying the underlying acoustic evidence. To address this issue, we first introduce spoofing evidence-grounded audio reasoning (SEAR), a four-task AQA benchmark to evaluate ALM-based ADD through acoustic evidence identification and quantification, deepfake detection, and forensic rationale generation. We further propose a bona-fide-based acoustic evidence agent (BAEA), which equips a frozen ALM with controlled acoustic tools under \\textsc{fixed} or \\textsc{adaptive} evidence-acquisition policies. Experiments with six ALMs reveal a clear gap between plausible rationales and verifiable acoustic evidence reasoning, while BAEA-\\textsc{Fixed} improves final verdicts and forensic rationales on both evaluation partitions. Controlled interventions further show that misleading evidence degrades both detection and grounding performance.


Key findings
ALMs can generate lexically plausible rationales but struggle to reliably identify or quantify corresponding acoustic evidence. The proposed BAEA-FIXED significantly improves deepfake detection (reduces EER) and acoustic evidence grounding (improves Ref-G) compared to vanilla ALMs. Misleading acoustic evidence substantially degrades both deepfake detection and rationale grounding performance.
Approach
The authors introduce SEAR, a benchmark with four tasks: deepfake verdict, forgery cue identification, acoustic feature measurement, and forensic rationale generation. They then propose BAEA, an agent that augments ALMs with a controlled signal-analysis tool and a bona-fide-based acoustic reference, operating under either a FIXED policy (measuring all features) or an ADAPTIVE policy (ALM selects features). This allows for evidence-grounded reasoning without updating the ALM's backbone parameters.
Datasets
ASVspoof 2019 LA (19LA), ASVspoof 2021 LA (21LA)
Model(s)
Qwen2-Audio-7B-Instruct, Qwen2.5-Omni-7B, MiniCPM-o-4.5, MOSS-Audio-8B-Instruct, Gemini-3.1-Flash-Lit, GPT-Audio-1.5
Author countries
United Kingdom, Singapore, China