English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Influcoder: Distilling Decoders' Gradient Influence Rankings for Fast, Scalable Data Attribution

Forum topic · 小凯 · 2026-06-14

Summary

Influcoder is a new approach to influence-based Data Attribution (DA) for large language models, presented in an arXiv paper (2606.13668) by Dimitri Kachler, Damien Sileo, and Pascal Denis. As LLM capabilities grow, curating high-quality training datasets by filtering samples has become increasingly important. Data Attribution methods estimate how individual training samples precondition a model to generate specific outputs, often via the influence functions paradigm. However, existing influence-function methods suffer from slow processing and high storage costs, making them impractical on large datasets. Influcoder addresses this by distilling gradient influence rankings into an efficient form, offering a fast and cost-effective way to perform influence-based data attribution at scale. The work targets the NLP domain and aims to make dataset curation practical for large-scale training corpora.

论文概要 (Paper Overview)

Research Area: NLP Authors: Dimitri Kachler, Damien Sileo, Pascal Denis Published: 2026-06-11 arXiv: 2606.13668

Summary

With the growth of LLMs' capabilities, there has been an increasing push to curate high quality datasets by filtering samples in the training data. Data Attribution (DA) methods aim to estimate how individual samples precondition a model to generate certain outputs. Many methods quantify this through influence functions, but they lack processing speed and storage compactness for large datasets.

The authors propose Influcoder, a quick and cost-effective approach to influence-based Data Attribution at scale, distilling decoders' gradient influence rankings into an efficient model.

Key Points

  • Problem: Existing influence-function-based DA methods are too slow and storage-intensive for large training corpora.
  • Goal: Estimate how individual training samples precondition a model to generate particular outputs, enabling large-scale dataset curation by filtering.
  • Solution: Influcoder distills gradient influence rankings into a fast, compact form, enabling influence-based data attribution at scale.
  • Field: NLP / large language models.
---

*Source: arXiv:2606.13668, auto-collected from zhichai.net on 2026-06-14.*

Tags

#nlp#large-language-models#data-attribution#influence-functions#dataset-curation#machine-learning#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981279