English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models

Forum topic · 小凯 · 2026-08-24

Summary

ConceptGuard is a new benchmark by Sahil Kale and Ian Harris (arXiv:2608.20338) that evaluates context-sensitive machine unlearning in large language models. The authors argue that effective unlearning must operate at the concept level: harmful applications of knowledge should be removed while benign, beneficial uses remain intact. To test this, they introduce the notion of dual-use concepts—knowledge that can be applied both harmfully and benignly—and build a benchmark whose forget and retain sets are explicitly complementary in concept usage. Evaluation is intent-sensitive and targets maximizing contextual separation to promote safer model behavior. Experiments show that current unlearning techniques perform poorly in this setting, exhibiting weak contextual separation, low ROUGE and concept-level metric scores, a strong unlearning-utility trade-off, and inconsistent concept-level control across methods. The paper highlights a gap in existing unlearning benchmarks, which rely on disjoint fact sets and simple fact recall, and positions ConceptGuard as a more meaningful standard for evaluating safe, concept-level knowledge removal in LLMs.

Paper Overview

  • Field: NLP
  • Authors: Sahil Kale, Ian Harris
  • Published: 2026-08-22
  • arXiv: 2608.20338
  • Abstract

    Large Language Models (LLMs) increasingly require selective removal of harmful or sensitive knowledge, called unlearning, yet existing methods and benchmarks fail to evaluate this capability completely. Current approaches rely on disjoint forget and retain sets composed of independent facts, and measure success with simple direct fact recall. This framing misses a key requirement of unlearning: the ability to eliminate harmful behaviors while preserving benign, beneficial knowledge.

    The authors argue that effective unlearning must operate at the concept level, ensuring full removal of unsafe applications while retaining their correct and useful uses—achieving a complete and meaningful unlearning in a conceptual sense.

    ConceptGuard Benchmark

  • Introduces the notion of dual-use concepts: concepts that can be used both harmfully and benignly.
  • The benchmark's forget and retain sets are explicitly complementary in concept usage.
  • Enables unlearning to be explored and evaluated at the concept level rather than at a sparse fact level.
  • Evaluation is intent-sensitive, aiming to maximize contextual separation to promote safer behavior.
  • Findings

  • Current unlearning techniques perform poorly under this setting, with weak contextual separation.
  • Both ROUGE and concept-level metrics show poor performance.
  • Results reveal a strong unlearning–utility trade-off.
  • There is limited improvement in context sensitivity and inconsistent concept-level control across methods.

Original Abstract (excerpt)

> Large Language Models (LLMs) increasingly require selective removal of harmful or sensitive knowledge, called unlearning, yet existing methods and benchmarks fail to evaluate this capability completely.

---

*Auto-collected on 2026-08-24*

Tags

#llm#machine-unlearning#benchmark#nlp#ai-safety#dual-use-concepts#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633913