Boundary and Intra-Segment Learning for Partial Audio Deepfake Localization

Authors: Zhe Ye, Xiangui Kang, Minhua Huang, Kai Wu, Kong Aik Lee, Chng Eng Siong

Published: 2026-09-22 07:49:37+00:00

AI Summary

This paper introduces Boundary and Intra-Segment Learning (BISL) for localizing partial audio deepfakes by modeling both authenticity transitions and internal segment characteristics. BISL uses boundary learning to differentiate authenticity changes from acoustic variations and intra-segment learning to capture and enhance consistency within continuous bona fide and spoofed segments. By integrating frame, boundary, and segment information, BISL significantly improves fine-grained deepfake localization performance.

Abstract

Partial audio deepfakes manipulate only selected speech regions, making them difficult to be localized. Existing methods exploit boundary cues for partial deepfake localization, but primarily focus on identifying boundary positions rather than modeling the feature changes that characterize authenticity transitions. Meanwhile, the internal characteristics of continuous bona fide and spoofed segments remain underexplored. In this paper, we propose Boundary and Intra-Segment Learning (BISL), which introduces boundary learning to model feature differences between adjacent frames and distinguish authenticity transitions from general acoustic variations. In addition, intra-segment learning captures the overall characteristics of continuous bona fide and spoofed segments while enhancing feature consistency within each segment. By jointly learning frame, boundary, and segment information, BISL enables more effective fine-grained partial audio deepfake localization. Experiments on multiple localization benchmarks show that BISL achieves an EER of 2.52\\% and an F1-score of 97.40\\% on PartialSpoof, outperforming the compared methods, while maintaining competitive performance on HAD and improved cross-dataset performance on LPS. The code will be made publicly available upon acceptance.


Key findings
BISL achieved an EER of 2.52% and an F1-score of 97.40% on PartialSpoof, outperforming compared methods. It also showed competitive performance on HAD and improved cross-dataset performance on LPS, demonstrating better generalization to unseen data distributions. Ablation studies confirmed that both boundary and intra-segment learning contribute to the overall effectiveness of the approach.
Approach
The proposed BISL approach jointly learns frame, boundary, and segment information. Boundary learning models feature differences between adjacent frames to identify authenticity transitions, while intra-segment learning captures the overall characteristics of continuous bona fide and spoofed segments and enhances feature consistency within them. This multi-task training objective leverages frame-level, segment-level, and boundary-level supervision.
Datasets
PartialSpoof (PS), Half-truth Audio Detection (HAD), LlamaPartialSpoof (LPS)
Model(s)
WavLM-Large, Conformer, two-layer MLP heads
Author countries
China, Hong Kong, Singapore