English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Simple Domain Generalization for Strong Pixel-Level Image Tampering Detection Under Modern VLMs

Forum topic · 小凯 · 2026-07-22

Summary

Modern vision-language models (VLMs) such as ChatGPT, Gemini, and Qwen-Image have dramatically advanced image generation and editing, making pixel-level image tampering detection increasingly important yet challenging under cross-model and out-of-distribution shifts. This paper studies domain generalization for pixel-level tampering localization across diverse VLM-generated manipulations. The authors propose a simple but effective training framework built on two practical strategies: (1) a balanced minibatch sampling scheme that strategically samples tampered and authentic images within each batch, preventing biased optimization toward manipulation artifacts or clean-image priors and avoiding training collapse; and (2) a late-injection strategy in which the detector is first trained on a large base dataset until stable convergence, then exposed to a small amount of support data from emerging VLM distributions, improving adaptability without overfitting to the new domain. On out-of-distribution VLMs including GPT-Images-2.0, Gemini-3.1, FLUX.2, and Seedream 4.5, the framework improves average gIoU and cIoU over the prior SOTA PIXAR by 26.1% and 26.8%, respectively. The work is available as arXiv preprint 2607.18230 (cs.CV, cs.AI).

Paper Overview

Field: Computer Vision (CV) Authors: Yi Tang, Xinyi Shang, Jiacheng Cui, Sondos Mahmoud Bsharat, Jiacheng Liu, et al. (15 authors in total) Posted: 2026-07-20 arXiv: 2607.18230 Categories: cs.CV, cs.AI

Abstract

Modern vision-language models (VLMs) have significantly enhanced image generation and editing capabilities, making pixel-level image tampering detection increasingly important — yet it remains challenging under cross-model and out-of-distribution (OOD) shifts. This work studies domain generalization for pixel-level tampering detection in modern VLMs (such as ChatGPT, Gemini, and Qwen-Image), aiming to learn tampering localization models that remain robust across diverse VLM-generated manipulation distributions.

The authors propose a simple yet effective domain generalization training framework based on two practical strategies:

1. Balanced minibatch sampling: a scheme that strategically samples tampered and authentic images within each minibatch, preventing biased optimization toward manipulation artifacts or clean-image priors and avoiding training collapse. 2. Late-injection strategy: the detector is first trained on a large base dataset until stable convergence, and is then exposed to a small amount of support data from emerging VLM distributions — improving adaptability without overfitting to the new domain.

Results

On OOD VLMs (GPT-Images-2.0, Gemini-3.1, FLUX.2, Seedream 4.5), the framework improves average gIoU and cIoU over the prior SOTA method PIXAR by 26.1% and 26.8%, respectively.

---

*Auto-collected on 2026-07-22.*

Tags

#computer-vision#image-forensics#tampering-detection#domain-generalization#vision-language-models#deep-learning#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178446996