论文概要 (Paper Overview)
Research Area: NLP Authors: Dimitri Kachler, Damien Sileo, Pascal Denis Published: 2026-06-11 arXiv: 2606.13668
Summary
With the growth of LLMs' capabilities, there has been an increasing push to curate high quality datasets by filtering samples in the training data. Data Attribution (DA) methods aim to estimate how individual samples precondition a model to generate certain outputs. Many methods quantify this through influence functions, but they lack processing speed and storage compactness for large datasets.
The authors propose Influcoder, a quick and cost-effective approach to influence-based Data Attribution at scale, distilling decoders' gradient influence rankings into an efficient model.
Key Points
- Problem: Existing influence-function-based DA methods are too slow and storage-intensive for large training corpora.
- Goal: Estimate how individual training samples precondition a model to generate particular outputs, enabling large-scale dataset curation by filtering.
- Solution: Influcoder distills gradient influence rankings into a fast, compact form, enabling influence-based data attribution at scale.
- Field: NLP / large language models.
*Source: arXiv:2606.13668, auto-collected from zhichai.net on 2026-06-14.*