Explainable Deepfake Detection with Feature-robust Augmentation and Evidence-grounded Explanation Optimization
Authors: Zhu Xu, Jiaqi Tang, Pokai Chen, Yuxin Peng, Yang Liu
Published: 2026-08-21 09:30:34+00:00
AI Summary
This paper proposes a framework for explainable deepfake detection that addresses vulnerabilities to image quality degradation and factually flawed explanations. It introduces Feature-robust Augmentation with supervised contrastive learning and a mean-teacher architecture for robust detection, and an evidence-grounded preference optimization for generating accurate and complete explanations. The approach won first place in the ACM Multimedia 2026 Explainable Deepfake Detection Challenge.
Abstract
Explainable deepfake detection extends binary classification by requiring models to not only predict authenticity but also provide interpretable justifications. This expanded scope is critical in practice, where users like forensic analysts need insight into the rationale behind the detection. Despite advancements, current approaches suffer from two critical deficiencies: (1)vulnerability to image quality degradation: detection accuracy plummets on low-quality samples, while naive augmentation strategies may induce feature drift and impair performance as diversity expands. (2) factually flawed explanations: explanation models may omit manipulation evidence or hallucinate irrelevant details, undermining interpretability. To address it, we propose a framework with two innovations. For robust deepfake detection, we introduce Feature-robust Augmentation, which comprises diversified degradation-aware augmentation strategies, and a supervised contrastive learning pattern paired with a mean-teacher architecture that stabilizes features against augmentations through consistency constraints. For explanation, we devise an evidence-grounded preference optimization process that guides model to prioritize genuine manipulation traces by learning from chosen-rejected explanation pairs, where rejected samples are constructed via evidence omission or irrelevant information injection. The proposed approach wins the first place in ACM Multimedia 2026 Explainable Deepfake Detection Challenge.The code is available at https://github.com/oceanflowlab/EDD.git.