ModalFidelity: Routing Modalities for Deepfake Detection on a Budget

Authors: Oguzhan Baser, Kaan Kale, Sriram Vishwanath, Sandeep Chinchali

Published: 2026-09-29 06:51:33+00:00

Comment: 5 pages, 4 figures, 1 table. Submitted to ICASSP 2027

AI Summary

ModalFidelity is a lightweight routing mechanism for deepfake detection that efficiently allocates computational resources. It previews each video window and, under a fixed compute budget, decides which audio or visual streams are worth analyzing by forensic detectors. This approach significantly reduces computational cost while maintaining high accuracy in detecting localized deepfakes.

Abstract

Deepfakes no longer need to fake a whole video. Generators that read the transcript now alter only the few seconds in which a video's meaning turns, so a forgery hides in a small, unknown fraction of the video. Yet detectors still read every one-second window of both the audio and image streams, spending nearly all of their compute where nothing was altered. We observe that deciding where to look is far cheaper than looking. We present ModalFidelity, a lightweight router that previews each window and decides, before any forensic detector runs, which stream is worth reading, under a hard compute budget it can never exceed. On AV-Deepfake1M, reading at most a fifth of the windows, it is more accurate than gating after the detectors at 15.9x less compute, and retains over 96% of the accuracy of an oracle that knows where every forgery lies.


Key findings
ModalFidelity achieves an accuracy comparable to a clairvoyant oracle (within 2.6 percentage points at every budget) while using significantly less compute (15.9x fewer GFLOPs than late MoEs). It demonstrates that judiciously choosing which modality to analyze at each window is crucial for efficient and accurate deepfake detection, especially for targeted forgeries that alter only small parts of a video.
Approach
The system employs a lightweight router (ModalFidelity) that uses a MobileNetV2 preview encoder and an LSTM to analyze cheap previews of each window. Based on this analysis, the router decides which modality (audio, image, or both) to send to more expensive forensic detectors, all while adhering to a predefined compute budget. This decision-making process is trained using a combination of dynamic programming for an oracle policy and self-critical policy gradient for the actual router.
Datasets
AV-Deepfake1M
Model(s)
MobileNetV2, LSTM, W2V2-AASIST, GenD, AVH-Align
Author countries
USA