[论文] LESSER: Post-Training Data Selection with Output-Layer Gradients

研究领域: ML 作者: Lyuxin David Zhang, Eric Wong, Surbhi Goel, Anton Xue 发布时间: 2026-10-02 arXiv: 2610.03702

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: ML 作者: Lyuxin David Zhang, Eric Wong, Surbhi Goel, Anton Xue 发布时间: 2026-10-02 arXiv: 2610.03702

中文摘要

大语言模型后训练数据的选择对下游性能有重大影响。基于梯度的数据选择是一种流行方法,通过训练数据梯度与小型验证集梯度的对齐程度来排序。然而,使用全参数梯度排序需要对每个样本进行昂贵的前向和反向传播,使得大规模候选池的计算不可行。这引发一个自然的问题:能否以极低成本近似全梯度特征?幸运的是,我们发现输出层梯度足以实现有效的数据选择,且仅需更便宜的前向传播。我们将其实现为 LESSER——一个即插即用的选择方法包装器,将特征提取的 FLOP 成本在 SFT 上降低 9.7 倍,在 RL 基准上降低 3.0 倍,同时在下游任务上保持与全梯度性能一致。实验发现,即使输出层梯度和全梯度对单个样本的排序不同,它们选择的批次仍具有对齐的梯度。

原文摘要

The choice of post-training data for large language models substantially affects downstream performance. Gradient-based data selection is a popular approach that ranks training data by how well their gradients align with those of a small validation set. However, ranking with full-parameter gradients requires an expensive backward pass on every sample, making computation intractable for large candidate pools. This raises a natural question: can we approximate full-gradient features at a fraction of the cost? Conveniently, we find that output-layer gradients suffice for effective data selection, yet require only the cheaper forward pass. We implement this as LESSER, a drop-in wrapper for selection methods that reduces the feature-extraction FLOP cost by \(9.7\times\) for SFT and \(3.0\times\) fo...


*自动采集于 2026-10-06*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens