English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

AEGIS: A Holistic Benchmark for Evaluating Forensic Analysis of AI-Generated Academic Images

Forum topic · 小凯 · 2026-05-02

Summary

AEGIS is a holistic benchmark introduced to evaluate forensic analysis of AI-generated academic images. It advances over prior benchmarks in three ways. First, it covers domain-specific complexity across seven academic categories with 39 fine-grained subtypes, revealing intrinsic forensic difficulty: even GPT-5.1 reaches only 48.80% overall performance, and expert models achieve limited localization accuracy (IoU of 30.09%). Second, it models four prevalent academic forgery strategies across 25 generative models, with 11 models yielding average forensic accuracy below 50%, showing that forensic techniques lag behind generative advances. Third, it jointly evaluates detection, inference, and localization, revealing complementary strengths across model families: multimodal large language models reach 84.74% accuracy on text artifact identification, while specialized detectors peak at 79.54% on binary authenticity detection. Evaluating 25 leading MLLMs, 9 expert models, and 1 unified multimodal understanding-and-generation model, AEGIS serves as a diagnostic testbed exposing fundamental limitations in academic image forensics.

Overview

Research area: Computer Vision (CV) Authors: Shilin Lu, Qinying Huang, Kai Wang et al. Published: 2026-04-30 arXiv: 2604.28177

The authors introduce AEGIS, a holistic benchmark for evaluating forensic analysis of AI-generated academic images (A holistic benchmark for Evaluating forensic analysis of AI-Generated academic ImageS).

Key Advances

1. Domain-Specific Complexity: AEGIS covers seven academic categories with 39 fine-grained subtypes, exposing intrinsic forensic difficulty. Even GPT-5.1 reaches only 48.80% overall performance, and expert models achieve limited localization accuracy (IoU of 30.09%).

2. Diverse Forgery Simulations: The benchmark models four prevalent academic forgery strategies across 25 generative models. Of these, 11 models yield average forensic accuracy below 50%, indicating that forensic techniques lag behind generative advances.

3. Multi-Dimensional Forensic Evaluation: AEGIS jointly evaluates detection, inference, and localization, revealing complementary strengths across model families. Multimodal large language models (MLLMs) reach 84.74% accuracy on text artifact identification, while expert detectors peak at 79.54% on binary authenticity detection.

Evaluation Scope and Findings

The benchmark evaluates 25 leading MLLMs, 9 expert models, and 1 unified multimodal understanding-and-generation model. Results show that AEGIS functions as a diagnostic testbed, exposing fundamental limitations in academic image forensics and highlighting the gap between current forensic capabilities and increasingly capable generative models.

---

*Auto-collected on 2026-05-02.*

Tags

#aegis#benchmark#image-forensics#ai-generated-images#academic-images#mllm#deepfake-detection#computer-vision

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619036