English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Class Activation Mapping in Explainable Computer Vision: A Method-Centric Survey

Forum topic · 小凯 · 2026-08-14

Summary

This arXiv paper (2508.03414) by AmirHossein Eshghi, Hamid Saadatfar, and Seyyed Ali Hoseini surveys class activation mapping (CAM), one of the most widely used visual explanation families in explainable AI. CAM methods convert internal model evidence into heatmaps highlighting image regions, channels, tokens, or patches that support a target class or concept. Since the original CAM formulation in 2016, the field has expanded far beyond global-average-pooling CNN classifiers, now covering gradient-based post-hoc explanations, gradient-free scoring and ablation methods, high-resolution upsampling, weakly supervised localization and segmentation, transformer token attribution, causal and debiased approaches, and foundation-model-era methods using CLIP, DINO, or SAM. The survey synthesizes 57 method-centric publications since 2016, proposing a taxonomy that separates methods by attribution mechanism, architectural dependency, and evaluation target. Key trends include a shift from explaining single class scores in single-resolution CNN layers toward comparative, multi-layer, probabilistic, token-aware, and foundation-model-aware explanations. The authors also highlight persistent fragmentation in evaluation, where faithfulness, localization, robustness, computational cost, and human trust are measured under differing protocols.

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: AmirHossein Eshghi, Hamid Saadatfar, Seyyed Ali Hoseini
  • Published: 2026-08-13
  • arXiv: 2508.03414
  • Summary

    Class activation mapping (CAM) is one of the most widely used visual explanation families in explainable AI. Its goal is intuitive: convert internal model evidence into heatmaps that highlight image regions, convolutional channels, tokens, or patches supporting a target class or concept.

    Since the first CAM formulation in 2016, the field has expanded far beyond global-average-pooling CNN classifiers. CAM-style methods now include:

  • Gradient-based post-hoc explanations
  • Gradient-free scoring and ablation methods
  • High-resolution upsampling
  • Weakly supervised localization and segmentation
  • Transformer token attribution
  • Causal and debiased approaches
  • Foundation-model-era methods using CLIP, DINO, SAM, or feature distribution comparison
This survey synthesizes 57 method-centric publications since 2016. The authors develop a taxonomy that separates methods by attribution mechanism, architectural dependency, and evaluation target, then review gradient-based CAM, recent and hybrid CAM-style methods, and model- or architecture-aware approaches.

Key Trends

The dominant trend is a shift from explaining a single class score in a single-resolution CNN layer toward comparative, multi-layer, probabilistic, token-aware, and foundation-model-aware explanations.

At the same time, evaluation remains fragmented: faithfulness, localization, robustness, computational cost, and human trust are often measured under differing protocols. The survey therefore highlights not only each method's contributions, but also the gaps it leaves behind and the gaps subsequent methods attempt to fill.

---

*Auto-collected on 2026-08-14*

Tags

#explainable-ai#computer-vision#class-activation-mapping#cam#survey#arxiv#transformers#foundation-models

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633457