English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models

Forum topic · 小凯 · 2026-08-22

Summary

ConceptGuard is a new benchmark from researchers Sahil Kale and Ian Harris (arXiv 2608.20338) designed to evaluate context-sensitive unlearning in large language models. The authors argue that existing unlearning methods and benchmarks rely on disjoint forget and retain sets made of independent facts, measuring success through simple fact recall. This framing overlooks a key requirement of unlearning: removing harmful behavior while preserving benign, beneficial knowledge. The paper proposes that unlearning must operate at the concept level, ensuring that unsafe applications of a concept are fully removed while its correct, useful uses remain intact. To test this, the authors introduce the notion of dual-use concepts—concepts that can be used both harmlessly and maliciously—and build ConceptGuard, where forget and retain sets are explicitly complementary in concept usage. Experiments show that current unlearning techniques perform poorly under this setting, revealing a strong forget-utility tradeoff and only limited gains in contextual sensitivity. The benchmark highlights a significant gap between current machine unlearning capabilities and the demands of safe, selective knowledge removal in LLMs.

Paper Overview

Field: NLP Authors: Sahil Kale, Ian Harris Published: 2026-08-22 arXiv: 2608.20338

Abstract

Large language models face growing demand for selective knowledge removal, but existing methods and benchmarks cannot fully evaluate this capability. Current approaches rely on disjoint forget and retain sets composed of independent facts, using simple fact recall to measure success.

This framing ignores a critical requirement of unlearning: eliminating harmful behavior while preserving benign, beneficial knowledge. The authors argue that unlearning must operate at the concept level—ensuring complete removal of unsafe applications of a concept while maintaining its correct, useful usage.

Key Contributions

  • Introduce the notion of dual-use concepts: concepts that can be used both harmfully and benignly.
  • Build ConceptGuard, a benchmark in which the forget set and retain set are explicitly complementary in concept usage.
  • Show experimentally that current unlearning techniques perform poorly in this setting.

Findings

Experiments reveal a strong forget-utility tradeoff and only limited improvements in contextual sensitivity among existing unlearning methods, exposing a substantial gap between current capabilities and the real requirements of safe selective knowledge removal in LLMs.

---

*Auto-collected on 2026-08-22*

Tags

#machine-unlearning#llm-safety#benchmark#nlp#dual-use-concepts#large-language-models

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633785