Complementary rPPG-Derived and Lip-Region Frequency Cues for Talking-Face Deepfake Detection
Authors: Othmane Harraq, Tamer Aldwairi
Published: 2026-09-13 03:52:54+00:00
AI Summary
This research investigates the effectiveness of complementary visual-only cues, specifically rPPG-derived waveforms and lip-region Discrete Cosine Transform (DCT) coefficients, for detecting talking-face deepfakes. The study demonstrates that a fusion of these cues outperforms unimodal baselines, though their individual strengths vary significantly across different deepfake generation methods and transferability scenarios. The authors highlight the method-dependent complementarity of these cues for deepfake detection.
Abstract
Talking-face (TF) deepfakes are detected unevenly by rPPG-based methods across generators. We study two lightweight visual-only cues, rPPG-derived waveforms extracted by RhythmFormer and lip-region discrete cosine transform (DCT) coefficients, on the seven TF methods of Celeb-DF++ under a subject-independent protocol. In-domain, lip-region DCT matches or exceeds the rPPG-derived 1D ResNet on every method except SadTalker, and Concat fusion reaches AUC 0.891 against 0.824 and 0.827 for the unimodal baselines. Under leave-one-generator-out evaluation the cues split: each transfers clearly better to three held-out methods, and IP-LAP is near chance for both. Concat averages 0.798 but falls below rPPG alone where DCT transfers poorly, so static fusion only partly exploits this complementarity. Lip-region DCT outperforms full-face DCT on six of seven methods. We treat the rPPG-derived signal as an empirical cue and do not claim it is cardiac in origin.