Look Before You Judge: Training-Free Region Mining for Grounded and Explainable Deepfake Detection

Authors: Chia-Ling Chen, Yu-Ting Ta, Jian-Yu Jiang-Lin, Tai-Ming Huang, Ling Lo, Po-Ching Chen, Yan-Tsung Wang, Pei-Heng Li, Ling Zou, Hong-Han Shuai, Wen-Huang Cheng

Published: 2026-09-28 16:19:22+00:00

AI Summary

This paper introduces "Look Before You Judge," a training-free framework for explainable deepfake detection using Multimodal Large Language Models (MLLMs). The framework identifies image-specific candidate evidence regions by contrasting the MLLM's decoder-to-visual attention between an original image and its Gaussian-blurred counterpart. These regions are then individually inspected, and local evidence is integrated with global context to render a final verdict.

Abstract

Multimodal large language models (MLLMs) can explain deepfake verdicts in natural language, but such explanations are not necessarily visually grounded in the visual evidence underlying the prediction. A model may describe plausible artifacts inferred from language priors rather than from image evidence. Existing grounding methods improve visual reliance through decoding or attention interventions, but they generally strengthen grounding over the entire image, making them ill-suited for forensic artifacts that are subtle, spatially localized, and image-dependent. We propose Look Before You Judge, a training-free framework that formulates explainable deepfake detection as a sequential evidence acquisition process. Instead of directly predicting image authenticity from holistic visual reasoning, our framework first identifies image-specific candidate evidence regions by contrasting the MLLM's decoder-to-visual attention between an original image and its Gaussian-blurred counterpart. The identified regions are then inspected individually, and the resulting local evidence is integrated with the global image context before reaching a final verdict. The framework operates without manipulation masks, external forensic models, or parameter updates, making it directly applicable to off-the-shelf MLLMs. Across five open-source MLLMs on TriDF and MMTD-Set, our framework improves detection accuracy by up to 12.8%, reduces CHAIR by up to 33.4% and hallucination rate by up to 21.3%, and outperforms representative training-free decoding and attention methods.


Key findings
The framework significantly improves detection accuracy (up to 12.8%), reduces explanation hallucination rates (up to 21.3%), and lowers CHAIR (up to 33.4%) across various MLLMs and datasets. It outperforms existing training-free visual grounding methods by focusing on spatially localized evidence, demonstrating that explicit evidence acquisition is crucial for trustworthy and explainable deepfake detection.
Approach
The framework operates in three stages: first, it mines a blur-sensitive visual prior by comparing MLLM attention maps of an original image and its blurred version to identify detail-sensitive regions. Second, these high-response tokens are grouped into spatially coherent local forensic regions. Finally, the MLLM inspects each identified region individually with attention steering, and the filtered local evidence is aggregated with the global image context to produce a final deepfake detection and explanation.
Datasets
TriDF, MMTD-Set
Model(s)
InternVL-3.5-8B, InternVL-3.5-14B, Qwen3-VL-8B-Instruct, Qwen3.5-9B, MiMo-VL-7B, GPT-5, Gemini 2.5-Pro, Claude Sonnet 4.5
Author countries
Taiwan