Geometric Inconsistency Localization in Multi-View Image Sets

Authors: Xander Staelens, Albéric Loos, Bert Ramlot, Hannes Mareen, Peter Lambert, Glenn Van Wallendael

Published: 2026-09-25 13:25:56+00:00

Comment: 8 pages, accepted at the Deepfake Forensics Workshop (DFF 2026) at ACM Multimedia 2026

AI Summary

This paper introduces DeformView, a wide-baseline multi-view (MV) dataset with pixel-level annotations for geometric inconsistencies in images. The authors propose DEFECt3R, a lightweight learning-based classifier that uses cross-view feature relationships to localize these geometric inconsistencies, outperforming existing consistency-scoring methods in forensic tasks. This work establishes a benchmark and baseline for geometric inconsistency localization in wide-baseline MV image pairs, highlighting its potential for multimedia forensics.

Abstract

Novel view synthesis (NVS) models can produce realistic new views of the same scene from different viewpoints. However, these generated views are not always geometrically consistent with one another. Multi-view (MV) consistency has shown promise as a tool for evaluating these NVS models. Its potential for multimedia forensics, however, remains largely unexplored, particularly for localizing geometric inconsistencies across wide-baseline image pairs. To enable research in this direction, we introduce DeformView, a wide-baseline MV dataset with pixel-level annotations of geometric inconsistencies. Using DeformView, we evaluate state-of-the-art MV consistency-scoring methods and show that approaches developed for NVS evaluation transfer poorly to the forensic task of geometric inconsistency localization. To address this limitation, we propose DEFECt3R, a lightweight learning-based classifier that uses cross-view feature relationships to localize geometric inconsistencies at the pixel level. By learning from explicit supervision, including hard negatives from geometrically consistent yet deformed views, DEFECt3R improves localization performance and substantially reduces false positives compared to existing consistency-scoring methods. Ablation experiments further show that both feature representations and correspondence quality contribute to localization performance. Overall, our findings demonstrate that MV geometric consistency is a promising yet underexplored signal for multimedia forensics and establish a benchmark and baseline for geometric inconsistency localization in wide-baseline MV image pairs. Code and dataset are available at https://github.com/IDLabMedia/DeformView-DEFECt3R


Key findings
DEFECt3R consistently outperforms existing methods like MEt3R and TSED in pixel-level and pair-level geometric inconsistency localization, significantly reducing false positives. The study also found that both feature representations (DINO and MASt3R) and correspondence quality are crucial for performance, with correspondence quality having a larger impact. These findings demonstrate the promise of multi-view geometric consistency as a signal for multimedia forensics, especially when utilizing learned localization models.
Approach
The authors introduce DeformView, a dataset with pixel-level annotations of geometric inconsistencies created by deforming 3D object meshes. They propose DEFECt3R, a learning-based classifier that builds upon MEt3R by extracting and upsampling DINO feature maps, then establishing pixel correspondences using MASt3R. DEFECt3R learns to classify geometric inconsistencies at the pixel level by analyzing differences between aligned DINO and MASt3R features through a lightweight MLP classifier, trained with explicit supervision including hard negatives.
Datasets
DeformView (proposed dataset), Google Scanned Objects
Model(s)
DEFECt3R (proposed), MEt3R, TSED, DINO, FeatUp, MASt3R
Author countries
Belgium