Summary
ConceptGuard is a new benchmark introduced by Sahil Kale and Ian Harris (arXiv:2608.20338) that tests whether machine unlearning methods in large language models can operate at the concept level rather than the level of isolated facts. Existing unlearning approaches rely on disjoint forget and retain sets made of independent facts, with success measured by simple factual recall. This framing overlooks a core requirement of unlearning: removing harmful behaviors while preserving benign, useful knowledge. The authors argue that unlearning must ensure complete removal of unsafe applications of a concept while retaining its correct, beneficial usage. To evaluate this, they introduce the notion of dual-use concepts—concepts that can be used both harmlessly and harmfully—and build ConceptGuard, where the forget and retain sets are explicitly complementary in concept usage. Experiments show that current unlearning techniques perform poorly in this setting, revealing a strong forget-utility trade-off and only limited gains in context sensitivity. ConceptGuard provides a more realistic standard for evaluating selective knowledge removal in LLMs.
Paper Overview
Field: NLP
Authors: Sahil Kale, Ian Harris
Published: 2026-08-22
arXiv: 2608.20338
Abstract
There is a growing demand for selective knowledge removal in large language models, but existing methods and benchmarks cannot fully evaluate this capability. Current approaches rely on disjoint forget and retain sets composed of independent facts, measuring success with simple factual recall.
This framing ignores a key requirement of unlearning: eliminating harmful behavior while preserving benign, beneficial knowledge. The paper argues that unlearning must operate at the concept level, ensuring complete removal of unsafe applications of a concept while maintaining its correct and useful usage.
Key Contributions
- Introduces dual-use concepts — concepts that can be used both harmlessly and benignly — as a testbed for context-sensitive unlearning.
- Presents ConceptGuard, a benchmark in which the forget and retain sets are explicitly complementary in their usage of these concepts.
Findings
Experiments show that current unlearning techniques perform poorly under this setting, revealing:
- A strong forget-utility trade-off: removing harmful usage degrades benign usage of the same concept.
- Only limited improvement in context sensitivity among existing methods.
ConceptGuard highlights the gap between fact-level unlearning and true concept-level knowledge removal, setting a more rigorous standard for future LLM unlearning research.
---
*Auto-collected on 2026-08-22*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178633806