Agentic Tool-Augmented Reasoning for Explainable Image Forgery Detection

Authors: Zhiya Tan, Jing Huang, Changtao Miao, Lin Tan, Xin Zhang, Weiwei Feng, Jianshu Li, Joey Tianyi Zhou

Published: 2026-09-30 06:05:16+00:00

Comment: Accepted at ACM Multimedia 2026 (Oral)

AI Summary

This paper introduces Agentic Tool-Augmented Reasoning (ATAR), a framework for explainable image forgery detection inspired by human forensic workflows. ATAR integrates 22 specialized forensic tools and employs a Dual-Stream Forensic Reasoning paradigm for autonomous detection, localization, and explanation of forgeries through multi-turn reasoning, coupled with a Forensics Curriculum Learning strategy for training. It achieves state-of-the-art results on various forgery detection benchmarks and produces more faithful and grounded explanations than existing MLLM-based approaches.

Abstract

Conventional image forgery detection methods produce binary scores or pixel-level masks without interpretable evidence, while recent multimodal large language model (MLLM)-based approaches generate post-hoc explanations of predetermined classification results rather than reasoning from evidence. Inspired by the forensic workflow of human judicial experts, we propose Agentic Tool-Augmented Reasoning (ATAR), a framework integrating 22 specialized forensic tools across seven complementary domains to autonomously detect, localize, and explain image forgeries through multi-turn reasoning. A Dual-Stream Forensic Reasoning paradigm combines a high-level semantic anomaly path, which magnifies suspicious regions for fine-grained inspection, with a low-level forgery artifact path, which invokes forensic tools to extract objective evidence. We further introduce Forensics Curriculum Learning: during General Experience SFT, an automated teacher-student mentoring pipeline synthesizes multi-turn tool-usage reasoning trajectories; during Forensic Scene RL, a Tool Prior Curriculum guides early tool exploration and progressively transfers control to the agent, while a Structured Evidence Reward provides fine-grained process-level supervision. Experiments across IMDL, Deepfake detection, DMDL, and AIGC detection show that ATAR achieves 78.5% average image-level F1 on six zero-shot IMDL benchmarks, surpassing the strongest MLLM baseline by 11.8 percentage points, and remains competitive with specialized detectors on other tasks while producing substantially more faithful and grounded explanations.


Key findings
ATAR significantly outperforms MLLM baselines, achieving 78.5% average image-level F1 on six zero-shot IMDL benchmarks, an 11.8 percentage point improvement over the strongest MLLM baseline. It remains competitive with specialized detectors across deepfake, DMDL, and AIGC detection tasks, while drastically reducing reasoning hallucination to 7.2% on true positives, demonstrating superior faithfulness and groundedness in its explanations.
Approach
ATAR leverages a Multimodal Large Language Model (MLLM) equipped with 22 forensic tools and a Dual-Stream Forensic Reasoning paradigm to mimic human forensic experts. It uses a high-level semantic anomaly path and a low-level forgery artifact path (invoking forensic tools) for evidence collection. The model is trained using Forensics Curriculum Learning, which includes General Experience SFT for synthesizing multi-turn reasoning trajectories and Forensic Scene RL with a Tool Prior Curriculum and Structured Evidence Reward for fine-grained supervision and tool selection.
Datasets
CASIA2, IMD2020, FantasticReality, NeXT-IMDL, AutoSplice, OpenForensics, DocTamper, CASIAv1+, Columbia, NIST16, Coverage, Korus, CocoGlide, T-SROIE, FSTS-1.5k, GenImage++
Model(s)
Qwen3-VL-8B-Instruct (fine-tuned with LoRA), Qwen3-VL-235B-A22B (as MLLM judge), SAM (Segment Anything Model), DINOv3
Author countries
Singapore, China