English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Concept-Based Abductive and Contrastive Explanations for Deep Neural Network Behaviors

Forum topic · 小凯 · 2026-05-10

Summary

This paper introduces concept-based abductive and contrastive explanations for deep neural networks, combining two research threads: concept-based explanations that describe predictions via human-understandable high-level concepts, and formal abductive/contrastive explanations that compute minimal causally relevant input features. Existing approaches either lack a causal link between concepts and predictions or can only infer explanations involving a single concept, while feature-based methods are limited to low-level features like pixels. The authors propose explanations that capture minimal sets of high-level concepts causally relevant to model outputs. They present a family of algorithms to enumerate all minimal explanations, using concept erasure procedures to establish causality. By aggregating such explanations, the method explains both individual image predictions and model behavior across image collections for user-specified shared behaviors. The approach was evaluated across multiple models, datasets, and behaviors, demonstrating its effectiveness in producing helpful, user-friendly explanations.

Paper Overview

Research Area: Machine Learning

Authors: Ronaldo Canizales, Divya Gopinath, Corina Păsăreanu, Ravi Mangal

Published: 2026-05-07

arXiv: 2605.06640

Abstract

Concept-based explanations offer a promising approach to explaining deep neural network predictions in terms of high-level, human-understandable concepts. However, existing methods either fail to establish a causal connection between concepts and model predictions, or have limited expressiveness, only able to infer causal explanations involving a single concept.

Meanwhile, parallel work on formal abductive and contrastive explanations computes minimal sets of input features that are causally relevant to model outcomes, but considers only low-level features such as pixels.

Merging these two lines of work, this paper introduces the notion of concept-based abductive and contrastive explanations, which capture minimal sets of high-level concepts causally relevant to model outcomes. The authors then propose a family of algorithms to enumerate all minimal explanations, leveraging concept erasure procedures to establish causality.

By appropriately aggregating such explanations, the approach enables understanding of model predictions not only on individual images, but also on collections of images exhibiting user-specified shared behaviors. The method was evaluated across multiple models, datasets, and behaviors, demonstrating its effectiveness in computing helpful, user-friendly explanations.

---

*Auto-collected on 2026-05-10*

Tags

#machine-learning#explainable-ai#neural-networks#formal-methods#interpretability#arxiv#deep-learning

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619701