English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Measuring Task-Agnostic Training Data Influence Across Language Model Pretraining

Forum topic · 小凯 · 2026-08-15

Summary

Researchers propose a task-agnostic method for measuring training data influence in language model pretraining. Instead of relying on downstream tasks or validation sets as attribution targets, an example's influence is defined by how much its gradient update reduces the squared distance to a pretraining run's final parameters, estimated from intermediate checkpoints without retraining. Applying the method to 18 configurations from the Pythia and PolyPythia suites, the authors find systematic temporal shifts in influential data: literature-related data align more strongly with the trajectory toward final parameters early in training, while STEM data become more strongly aligned in later stages. This qualitative crossover is broadly consistent across model configurations. The approach offers a tractable, trajectory-level view of how influential data evolve throughout pretraining, complementing influence analyses tied to specific tasks or validation sets. Paper: arXiv 2608.13515.

Paper Overview

Field: NLP Authors: Yuto Nishida, Hirokazu Kiyomaru, Yusuke Oda, Takashi Kodama, Chaoran Liu, Daisuke Kawahara, Yusuke Miyao, Max Müller-Eberstein, Masaru Isonuma arXiv: 2608.13515

Abstract

Measuring training data influence consistently across language model pretraining is challenging. It is difficult to select downstream tasks or validation sets representative of a model's general capabilities, and reliance on task performance at intermediate checkpoints complicates comparisons across training. The authors propose a measure of training data influence that does not require selecting a downstream task or validation set as the attribution target. Specifically, an example's influence is defined by how much its gradient update reduces the squared distance to the final parameters of a given pretraining run, and this quantity is estimated from intermediate checkpoints without retraining.

Key Findings

  • Applying the method to 18 configurations from the Pythia and PolyPythia suites reveals systematic temporal changes in influential data.
  • Early in training, literature-related data are more strongly aligned with the trajectory toward the final parameters.
  • STEM data become more strongly aligned in later training stages.
  • This qualitative crossover is broadly consistent across model configurations.
  • The results provide a tractable trajectory-level view of how influential data change throughout pretraining, complementing influence analyses defined with respect to specific downstream tasks or validation sets.

Tags

#nlp#language-models#training-data-influence#pretraining#pythia#data-attribution#arxiv#machine-learning

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633511