Complementary rPPG-Derived and Lip-Region Frequency Cues for Talking-Face Deepfake Detection

Authors: Othmane Harraq, Tamer Aldwairi

Published: 2026-09-13 03:52:54+00:00

AI Summary

This research investigates the effectiveness of complementary visual-only cues, specifically rPPG-derived waveforms and lip-region Discrete Cosine Transform (DCT) coefficients, for detecting talking-face deepfakes. The study demonstrates that a fusion of these cues outperforms unimodal baselines, though their individual strengths vary significantly across different deepfake generation methods and transferability scenarios. The authors highlight the method-dependent complementarity of these cues for deepfake detection.

Abstract

Talking-face (TF) deepfakes are detected unevenly by rPPG-based methods across generators. We study two lightweight visual-only cues, rPPG-derived waveforms extracted by RhythmFormer and lip-region discrete cosine transform (DCT) coefficients, on the seven TF methods of Celeb-DF++ under a subject-independent protocol. In-domain, lip-region DCT matches or exceeds the rPPG-derived 1D ResNet on every method except SadTalker, and Concat fusion reaches AUC 0.891 against 0.824 and 0.827 for the unimodal baselines. Under leave-one-generator-out evaluation the cues split: each transfers clearly better to three held-out methods, and IP-LAP is near chance for both. Concat averages 0.798 but falls below rPPG alone where DCT transfers poorly, so static fusion only partly exploits this complementarity. Lip-region DCT outperforms full-face DCT on six of seven methods. We treat the rPPG-derived signal as an empirical cue and do not claim it is cardiac in origin.


Key findings
Lip-region DCT often matches or exceeds rPPG-derived cues, except for SadTalker deepfakes, and Concat fusion significantly improves overall detection (AUC 0.891). Under leave-one-generator-out evaluation, the cues exhibit strong complementarity, with each transferring better to specific held-out methods. Lip-region DCT consistently outperforms full-face DCT for deepfake detection.
Approach
The authors extract two lightweight visual-only cues: rPPG-derived waveforms using RhythmFormer and lip-region DCT coefficients. These cues are then processed by separate encoders (1D ResNet for rPPG, MLP for DCT) and combined using various fusion architectures, with 'Concat fusion' proving most effective. The system is evaluated under subject-independent and leave-one-generator-out protocols.
Datasets
Celeb-DF++ (talking-face subset)
Model(s)
RhythmFormer, 1D ResNet, MLP, Gated fusion, Concat fusion, Cross-attention fusion
Author countries
USA