Interpretable Deepfake Detection in Videos via Explicit Forensic Features and Temporal Modeling

Authors: Chahira Benhama, Mohand Saïd Allili, Assia Hamadene

Published: 2026-10-02 14:34:24+00:00

Comment: 10

AI Summary

This paper introduces an interpretable deepfake detection framework for videos that explicitly models spatially and temporally coherent facial features. It transforms videos into identity-consistent facial trajectories, representing each frame with 68 structured descriptors across photometric, textural, geometric, and compression-based domains. These multi-domain features are then processed by an LSTM network to capture temporal dependencies and subtle irregularities for robust deepfake detection and improved multi-dataset generalization.

Abstract

Deepfake detection in videos remains challenging, as manipulated content may appear visually consistent at the frame level while exhibiting subtle temporal inconsistencies. This paper introduces an interpretable deepfake detection framework that models spatially and temporally coherent facial features in video sequences. Unlike end-to-end deep models relying on implicit representations, the proposed approach explicitly encodes physically grounded forensic cues, enabling transparent analysis and improved multi-dataset generalization. The pipeline transforms videos into identity-consistent facial trajectories, segments them into fixed-length temporal windows, and represents each frame using 68 structured descriptors spanning four complementary domains: photometric, textural, geometric, and compression-based features. These descriptors provide a compact multi-domain representation of manipulation artifacts and are processed by a Long Short-Term Memory (LSTM) network to capture temporal dependencies and subtle irregularities. Evaluation on four benchmark datasets, FaceForensics++, Celeb-DF v2, a curated subset of the DeepFake Detection Challenge (DFDC), and DeeperForensics, yields strong and consistent F1-scores of 98.0%, 91.0%, 97.6%, and 96.2%, respectively. The approach also demonstrated a good cross-dataset generalization, providing a robust and interpretable solution for video deepfake detection.


Key findings
The framework achieved strong and consistent F1-scores of 98.0% on FaceForensics++, 91.0% on Celeb-DF v2, 97.6% on DFDC subset, and 96.2% on DeeperForensics. It also demonstrated good cross-dataset generalization with 93.8% F1-score on ForgeryNet and 89.5% on WildDeepfake without retraining. The ablation study confirmed that combining all four feature families (compression, photometric, textural, geometric) consistently yielded the best performance, with compression and photometric features having the largest impact.
Approach
The approach extracts identity-consistent facial trajectories from videos and segments them into fixed-length temporal windows. Each frame within these windows is represented by 68 explicit, physically grounded forensic features spanning photometric, textural, geometric, and compression domains. A Long Short-Term Memory (LSTM) network then processes these structured feature sequences to identify temporal inconsistencies indicative of deepfakes.
Datasets
FaceForensics++, Celeb-DF v2, DeepFake Detection Challenge (DFDC) subset, DeeperForensics, ForgeryNet, WildDeepfake
Model(s)
Long Short-Term Memory (LSTM) network
Author countries
Canada