Synthetic speech detection in Brazilian Portuguese through accent-related features
Authors: Pedro H. L. Leite, Pedro Benevenuto Valadares, Luiz Wagner Pereira Biscainho
Published: 2026-09-20 18:52:40+00:00
Comment: \\c{opyright} 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
AI Summary
This work introduces a speech deepfake detection methodology for Brazilian Portuguese (pt-BR) by exploiting the dialectal inconsistencies of synthetic speech. It combines multilingual phone recognizers and classical signal processing to extract phoneme-level features with high geographic variance. The approach reveals that the distributional gap in these accent-related features effectively distinguishes natural and synthetic voices, establishing dialectal inconsistency as an interpretable cue for spoofing detection.
Abstract
Leading commercial and open-source Text-to-Speech (TTS) models fail to emulate the regional phonetic diversity of Brazilian Portuguese (pt-BR). By aggregating disparate dialects into a single training distribution, they generate a synthetic diluted accent: a phonetic profile attempting to represent all regional distributions simultaneously, but ultimately carrying phonological ambiguity dissociated from natural socio-phonetic realizations. This work introduces a speech deepfake detection methodology combining multilingual phone recognizers with classical signal processing to extract phoneme-level features in consonantal and vocalic realizations with high geographic variance. The analysis reveals that the distributional gap over these features suffices to distinguish natural and synthetic voices through unsupervised Kernel Density Estimation, establishing dialectal inconsistency as a useful and interpretable feature for spoofing detection in pt-BR. Evaluation on pt-BR anti-spoofing datasets shows that these explainable, lightweight, low-dimensional features can boost the performance of foundation models on the task, and show generalization capabilities in a cross-dataset leave-one-out setup.