Agentic Tool-Augmented Reasoning for Explainable Image Forgery Detection
Authors: Zhiya Tan, Jing Huang, Changtao Miao, Lin Tan, Xin Zhang, Weiwei Feng, Jianshu Li, Joey Tianyi Zhou
Published: 2026-09-30 06:05:16+00:00
Comment: Accepted at ACM Multimedia 2026 (Oral)
AI Summary
This paper introduces Agentic Tool-Augmented Reasoning (ATAR), a framework for explainable image forgery detection inspired by human forensic workflows. ATAR integrates 22 specialized forensic tools and employs a Dual-Stream Forensic Reasoning paradigm for autonomous detection, localization, and explanation of forgeries through multi-turn reasoning, coupled with a Forensics Curriculum Learning strategy for training. It achieves state-of-the-art results on various forgery detection benchmarks and produces more faithful and grounded explanations than existing MLLM-based approaches.
Abstract
Conventional image forgery detection methods produce binary scores or pixel-level masks without interpretable evidence, while recent multimodal large language model (MLLM)-based approaches generate post-hoc explanations of predetermined classification results rather than reasoning from evidence. Inspired by the forensic workflow of human judicial experts, we propose Agentic Tool-Augmented Reasoning (ATAR), a framework integrating 22 specialized forensic tools across seven complementary domains to autonomously detect, localize, and explain image forgeries through multi-turn reasoning. A Dual-Stream Forensic Reasoning paradigm combines a high-level semantic anomaly path, which magnifies suspicious regions for fine-grained inspection, with a low-level forgery artifact path, which invokes forensic tools to extract objective evidence. We further introduce Forensics Curriculum Learning: during General Experience SFT, an automated teacher-student mentoring pipeline synthesizes multi-turn tool-usage reasoning trajectories; during Forensic Scene RL, a Tool Prior Curriculum guides early tool exploration and progressively transfers control to the agent, while a Structured Evidence Reward provides fine-grained process-level supervision. Experiments across IMDL, Deepfake detection, DMDL, and AIGC detection show that ATAR achieves 78.5% average image-level F1 on six zero-shot IMDL benchmarks, surpassing the strongest MLLM baseline by 11.8 percentage points, and remains competitive with specialized detectors on other tasks while producing substantially more faithful and grounded explanations.