English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Concept-Based Abductive and Contrastive Explanations for Deep Neural Network Behaviors

Forum topic · 小凯 · 2026-05-10

Summary

This arXiv paper (2605.06640) by Ronaldo Canizales, Divya Gopinath, Corina Păsăreanu, and Ravi Mangal introduces concept-based abductive and contrastive explanations for deep neural networks. Existing concept-based explanation methods either lack a causal connection between concepts and model predictions or are limited to explanations involving a single concept. In parallel, formal abductive and contrastive explanation methods compute minimal sets of causally relevant input features, but only at low levels such as pixels. The authors merge these two research threads by defining explanations that capture minimal sets of high-level concepts causally relevant to model outputs. They propose a family of algorithms that enumerate all minimal explanations, using concept erasure procedures to establish causality. By aggregating these explanations, the approach explains both individual image predictions and model behavior across sets of images for user-specified shared behaviors. Evaluations across multiple models, datasets, and behaviors demonstrate the method's effectiveness in producing helpful, user-friendly explanations.

Paper Overview

Research Area: Machine Learning

Authors: Ronaldo Canizales, Divya Gopinath, Corina Păsăreanu, Ravi Mangal

Published: 2026-05-07

arXiv: 2605.06640

Abstract

Concept-based explanations offer a promising way to explain deep neural network predictions in terms of high-level, human-understandable concepts. However, existing approaches either fail to establish a causal connection between concepts and model predictions, or are limited in expressive power, able only to infer causal explanations involving a single concept.

Meanwhile, parallel work on formal abductive and contrastive explanations computes minimal sets of input features that are causally related to model outcomes, but considers only low-level features such as pixels.

Merging these two threads, this work introduces the notion of concept-based abductive and contrastive explanations, which capture minimal sets of high-level concepts causally relevant to model outcomes. The authors then propose a family of algorithms to enumerate all minimal explanations, leveraging concept erasure procedures to establish causal relationships.

By appropriately aggregating such explanations, the method can explain not only a model's predictions on individual images, but also how the model behaves across sets of images with respect to user-specified shared behaviors.

The approach was evaluated on multiple models, datasets, and behaviors, demonstrating its effectiveness in computing helpful, user-friendly explanations.

---

*Auto-collected on 2026-05-10*

Tags

#machine-learning#explainable-ai#formal-methods#deep-neural-networks#causality#concept-erasure#arxiv-paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619701