ModalFidelity: Routing Modalities for Deepfake Detection on a Budget
Authors: Oguzhan Baser, Kaan Kale, Sriram Vishwanath, Sandeep Chinchali
Published: 2026-09-29 06:51:33+00:00
Comment: 5 pages, 4 figures, 1 table. Submitted to ICASSP 2027
AI Summary
ModalFidelity is a lightweight routing mechanism for deepfake detection that efficiently allocates computational resources. It previews each video window and, under a fixed compute budget, decides which audio or visual streams are worth analyzing by forensic detectors. This approach significantly reduces computational cost while maintaining high accuracy in detecting localized deepfakes.
Abstract
Deepfakes no longer need to fake a whole video. Generators that read the transcript now alter only the few seconds in which a video's meaning turns, so a forgery hides in a small, unknown fraction of the video. Yet detectors still read every one-second window of both the audio and image streams, spending nearly all of their compute where nothing was altered. We observe that deciding where to look is far cheaper than looking. We present ModalFidelity, a lightweight router that previews each window and decides, before any forensic detector runs, which stream is worth reading, under a hard compute budget it can never exceed. On AV-Deepfake1M, reading at most a fifth of the windows, it is more accurate than gating after the detectors at 15.9x less compute, and retains over 96% of the accuracy of an oracle that knows where every forgery lies.