Cross-modal Translation via Conditional Latent Denoising for Video Deepfake Detection
Authors: Xinzhe Li, Youzhi Tu, Kong Aik Lee
Published: 2026-09-27 09:16:15+00:00
AI Summary
This paper introduces Cross-modal Translation via Conditional Latent Denoising (CTCLD), a framework for video deepfake detection that leverages cross-modal correspondences between audio and visual signals. By connecting heterogeneous modalities in latent spaces and performing bidirectional latent denoising, CTCLD effectively captures subtle inconsistencies characteristic of manipulated signals. The method establishes a Bayesian foundation by decomposing the audio-visual joint distribution to enable smooth cross-domain information transfer and improve detection performance.
Abstract
The growing threat of video deepfakes necessitates multimodal detection. Beyond serving as independent indicators of authenticity, audio and visual signals have intrinsic dependencies that also provide an essential criterion for detection. Previous methods often overlook the cross-modal correspondences, hindering information transfer between domains and leaving crucial detection cues unexplored. To address this challenge, we propose a framework called Cross-modal Translation via Conditional Latent Denoising (CTCLD) for video deepfake detection. It connects the distinct distributions of heterogeneous modalities in latent spaces, enabling smooth cross-domain information transfer to improve detection performance. We first establish a Bayesian foundation by decomposing the audio-visual joint distribution. Subsequently, CTCLD translates both modalities via bidirectional latent denoising conditioned on each other, effectively capturing subtle inconsistencies in the manipulated signals. Experimental results demonstrate that the proposed CTCLD enables comprehensive domain alignment, resulting in a robust video deepfake detection approach with competitive performance.