What I See is What I Hear: Deepfake Detection Across Diverse Hearing Abilities

Authors: Magdalena Pasternak, Malvika Jadhav, Palavi V. Bhole, Aviva Smith, Elaina Trapatsos, Vincent Bindschaedler, Roshan Peiris, Ersin Uzun, Patrick Traynor, Matthew Wright, Kevin R. B. Butler

Published: 2026-09-23 18:04:51+00:00

Comment: Proceedings of the Network and Distributed System Security (NDSS) Symposium 2027

AI Summary

This study investigates how human perception of audiovisual deepfakes varies across different hearing abilities. Through an in-person, mixed-methods study with 80 participants (hearing, hard-of-hearing, d/Deaf, and cochlear implant users), the research found that d/Deaf and hard-of-hearing (DHH) individuals were generally less accurate in detecting deepfakes than hearing persons (HPs), primarily due to misclassifying authentic clips as manipulated. The study highlights that deepfake detection heavily depends on the manipulated channel and the viewer's sensory access, emphasizing the need for accessible and tailored defenses for all users.

Abstract

The proliferation of audiovisual deepfakes has lowered the cost of fraud, impersonation, and misinformation, but their success ultimately depends on human perception. Detection requires integrating auditory and visual cues, yet security and privacy research has largely overlooked d/Deaf and hard-of-hearing (DHH) populations. We address this gap with an in-person, mixed-methods study of 80 participants: 31 hearing persons (HPs), 15 hard-of-hearing (HoH) participants, 17 d/Deaf participants, and 17 cochlear implant (CI) users. Each participant judged the authenticity of 30 clips, where manipulations spanned text-to-speech, voice conversion, lip-sync, or face-swap. DHH participants were less accurate than HPs overall (76.4% vs. 88.0%, p<.001), primarily because they more often classified authentic clips as manipulated (FPR: 29.7% vs. 11.2%). Differences depended strongly on the manipulated channel. For audio-only manipulations, HoH participants matched HPs (90.0% vs. 90.3%), followed by CI users (79.4%) and d/Deaf participants (41.2%). When clips contained an audiovisual manipulation, accuracy clustered between 84% and 87%, although performance still varied by manipulation method. Our work systematically characterizes how deepfakes affect DHH populations, highlighting the asymmetric risks audiovisual manipulations may pose to groups with different hearing abilities and the need for accessible, tailored defenses that support all users.


Key findings
DHH participants were less accurate overall (76.4%) compared to HPs (88.0%), largely due to higher false positive rates (29.7% vs. 11.2%). Detection accuracy varied significantly based on the manipulated channel: for audio-only manipulations, d/Deaf participants performed significantly worse (41.2%) than HPs (90.3%), while for visual-only and audiovisual manipulations, accuracies converged across groups. The study also found that DHH participants used a broader range of cues and often exhibited excessive skepticism, rejecting authentic clips more frequently.
Approach
The researchers conducted an in-person, mixed-methods study where 80 participants (divided into HPs, HoH, d/Deaf, and CI users) judged the authenticity of 30 audiovisual clips. These clips included authentic content and various manipulations: text-to-speech, voice conversion, lip-sync, or face-swap. Participants reported their judgment, confidence, and the cues influencing their decisions.
Datasets
High-Definition Talking-Face (HDTF) dataset
Model(s)
FaceFusion with InSwapper-128-fp16 and GPEN-BFR-2048 (for face-swap), SyncTalk (for lip-sync), HierSpeech++ (for text-to-speech), Seed-VC (for voice conversion)
Author countries
United States