MedForge-RSI: Medical Deepfake Detection via Recursive Self-Improvement

Authors: Zhihui Chen, Mengling Feng

Published: 2026-09-29 02:37:09+00:00

AI Summary

MedForge-RSI is a recursive self-improvement framework designed to enhance medical deepfake detectors in deployment settings without retraining their model weights. It enables a deployed detector to autonomously adapt by analyzing errors, accumulating experience, and developing image-analysis tools, with all improvements validated by an independent acceptance test. This framework significantly boosts detection accuracy and robustness to distribution shifts, particularly improving authentic-image recall.

Abstract

Text-guided image editors can generate high-fidelity medical deepfakes, challenging the reliability of clinical imagery. Although reasoning-based detectors perform strongly in distribution, they degrade substantially under deployment shift. MedForge-Reasoner, an 8B vision-language model trained with supervised fine-tuning and reinforcement learning, achieves 99.2% accuracy on its target distribution, yet misclassifies 40% of authentic scans, reaches only 77% accuracy on unseen generators, and falls to 59% under transmission distortion. Adapting such models through conventional retraining is costly, requiring large-scale supervision and expert-designed guidelines. We introduce MedForge-RSI, a recursive self-improvement framework that enables a deployed detector to adapt while keeping its model weights frozen. Over 20 rounds, the detector analyzes verified errors, accumulates reusable experience, and autonomously develops image-analysis tools, while an independent acceptance test retains only validated improvements. Across 49 registered configurations, MedForge-RSI increases average accuracy over four test sets from 75.0% to 84.4%. On a held-out 4,000-image evaluation, it improves clean accuracy from 76.5% to 87.9% and transmission-distorted accuracy from 59.0% to 70.5%, with the largest gains in authentic-image recall. Controlled analysis across all 49 trajectories identifies which self-improvement mechanisms replicate across seeds, shows acceptance testing to be the largest individual contributor, and reveals a taxonomy of failed adaptations. We release the complete trajectories, including all rejected and rolled-back changes.


Key findings
MedForge-RSI improved the average accuracy of the deployed detector across four test sets from 75.0% to 84.4%, and clean accuracy on a 4,000-image evaluation from 76.5% to 87.9%. The largest gains were observed in authentic-image recall and robustness to transmission distortion, increasing from 59.0% to 70.5%. The study identified acceptance testing as the largest contributor to improvements, and characterized several failure modes and safeguards for autonomous adaptation.
Approach
MedForge-RSI employs a recursive self-improvement loop where a frozen vision-language model, MedForge-Reasoner, reflects on its verified errors. It then proposes changes to two knowledge stores: an experience memory (structured entries with imaging context, visual cues, and cautions) and a toolbox (self-written image-analysis tools). Every proposed change undergoes a two-stage acceptance test, which includes a replay on the learning batch and an A/B comparison on an independent validation split, ensuring only validated improvements are adopted while core model weights remain fixed.
Datasets
MedForge-90K
Model(s)
MedForge-Reasoner (based on Qwen3-VL-8B-Instruct)
Author countries
Singapore