English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

LACUNA: A Testbed for Evaluating Localization Precision in LLM Unlearning

Forum topic · 小凯 · 2026-07-04

Summary

Large language models memorize sensitive training data, including personally identifiable information (PII), creating demand for reliable post hoc removal. State-of-the-art unlearning methods often follow a localize-first, unlearn-second paradigm, but existing benchmarks evaluate unlearning only at the output level, leaving unclear whether knowledge is truly erased from parameters or merely obfuscated. LACUNA, introduced by researchers including Matteo Boglioni, Thibault Rousset, and Siva Reddy (arXiv:2507.00481), is the first unlearning testbed with ground-truth parameter-level localization. It injects PII of synthetic individuals into predefined parameters of 1B and 7B OLMo base models via masked continual pretraining, enabling direct evaluation of whether unlearning methods target the weights responsible for knowledge storage. Benchmarking current SOTA methods shows that, despite strong output-level performance, existing approaches are highly imprecise and vulnerable to resurfacing attacks. When localization is accurate, even simple gradient-based unlearning achieves strong erasure and robustness, underscoring the importance of precise localization. LACUNA is released to complement behavioral evaluation.

Paper Overview

  • Field: NLP
  • Authors: Matteo Boglioni, Thibault Rousset, Siva Reddy
  • arXiv: 2507.00481
  • Key Points

  • LLMs memorize sensitive training data, including personally identifiable information (PII), creating a pressing need for reliable post hoc removal methods.
  • Unlearning has emerged as a promising solution; SOTA methods often follow a localize-first, unlearn-second paradigm that targets specific model parameters.
  • Existing benchmarks evaluate unlearning solely at the output level, leaving open whether unlearning truly erases knowledge from parameters or merely obfuscates it — a concern reinforced by successful resurfacing attacks.
  • What LACUNA Is

  • LACUNA is the first unlearning testbed with ground-truth parameter-level localization.
  • It injects PII of synthetic individuals into predefined parameters of 1B and 7B OLMo base models via masked continual pretraining.
  • This enables direct evaluation of whether unlearning methods target the weights responsible for storing knowledge.
  • Findings

  • Despite strong output-level performance, current SOTA unlearning methods are highly imprecise and vulnerable to resurfacing attacks.
  • When localization is accurate, even simple gradient-based unlearning achieves strong erasure and robustness against resurfacing attacks, highlighting the importance of precise localization.

Contribution

The authors release LACUNA to complement behavioral evaluation and drive further development of localization-based robust unlearning.

*Auto-collected on 2026-07-04.*

Tags

#llm-unlearning#machine-unlearning#privacy#pii#nlp#benchmark#arxiv#olmo

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208393