Cross-modal Translation via Conditional Latent Denoising for Video Deepfake Detection

Authors: Xinzhe Li, Youzhi Tu, Kong Aik Lee

Published: 2026-09-27 09:16:15+00:00

AI Summary

This paper introduces Cross-modal Translation via Conditional Latent Denoising (CTCLD), a framework for video deepfake detection that leverages cross-modal correspondences between audio and visual signals. By connecting heterogeneous modalities in latent spaces and performing bidirectional latent denoising, CTCLD effectively captures subtle inconsistencies characteristic of manipulated signals. The method establishes a Bayesian foundation by decomposing the audio-visual joint distribution to enable smooth cross-domain information transfer and improve detection performance.

Abstract

The growing threat of video deepfakes necessitates multimodal detection. Beyond serving as independent indicators of authenticity, audio and visual signals have intrinsic dependencies that also provide an essential criterion for detection. Previous methods often overlook the cross-modal correspondences, hindering information transfer between domains and leaving crucial detection cues unexplored. To address this challenge, we propose a framework called Cross-modal Translation via Conditional Latent Denoising (CTCLD) for video deepfake detection. It connects the distinct distributions of heterogeneous modalities in latent spaces, enabling smooth cross-domain information transfer to improve detection performance. We first establish a Bayesian foundation by decomposing the audio-visual joint distribution. Subsequently, CTCLD translates both modalities via bidirectional latent denoising conditioned on each other, effectively capturing subtle inconsistencies in the manipulated signals. Experimental results demonstrate that the proposed CTCLD enables comprehensive domain alignment, resulting in a robust video deepfake detection approach with competitive performance.


Key findings
CTCLD achieves competitive performance in video deepfake detection, outperforming existing unimodal and multimodal methods. It demonstrates high accuracy and AUC scores on both FakeAVCeleb (99.4% Acc, 99.5% AUC) and DFDC (98.5% Acc, 99.9% AUC) datasets. Ablation studies confirm the necessity of bidirectional cross-modal translation for robust detection, and visualization shows clear separation between real and fake video representations.
Approach
CTCLD first extracts latent embeddings for audio and visual modalities. It then uses a bidirectional latent denoising process, where each modality is translated and reconstructed conditioned on the other, to model cross-modal conditional distributions. A detection head then uses the concatenated enriched representations to identify forgery artifacts based on inconsistencies between the modalities.
Datasets
VoxCeleb2, FakeAVCeleb, DeepFake Detection Challenge (DFDC)
Model(s)
Diffusion Transformers (specifically, diffusion transformer blocks are used for denoising networks), encoders for audio and visual signals, a detection head.
Author countries
Hong Kong SAR